Setting expectations for anyone about to try this on a mobile workstation. Sustained GPU inference on battery is limited by the power the platform will deliver on battery, which is substantially less than on AC, and by the thermal envelope of a machine that is often on a soft surface.
The consequence is that the same model at the same settings will generate more slowly on battery, and the machine will get there by throttling rather than by failing. If you benchmark on AC and then use it on battery you will conclude something has broken.
Always state the power source when reporting local model throughput. A figure without it is not comparable to anything.