what do you use "the most powerful language model for its size" for?
* as in matches the token that the larger model would output
They would need the same vocabulary, etc. What else?
Even running a quantized and optimized LLM on a smartphone would kill battery life at minimum.
In the future(?), they will probably use the AI blocks instead of the GPU, which are very low power.