There are basically no useful models that run on phone hardware.
> Results vary by model size and quantization.
I bet they do.
Look, if you cant run models on your desktop, theres no way in hell they run on your phone.
The problem with all of these self hosting solutions is that the actual models you can run on them aren't any good.
Not like, “chat gpt a year ago” not good.
Like, “its a potato pop pop” no good.
Unsloth has a good guide on running qwen3 (1), and the tldr is basically, its not really good unless you run a big version.
The iphone 17 pro has 12GB of ram.
That is, to be fair, enough to run some small stable diffusion models, but it isnt enough to run run a decent quant of qwen3.
You need about 64 GB for that.
So… i dunno. This feels like a bunch of empty promises; yes, technically it can run some models, but how useful is it actually?
Self hosting needs next gen hardware.
This gen of desktop hardware isnt good enough, even remotely, to compare to server api options.
Running on mobile devices is probably still a way away.
(1) - https://unsloth.ai/docs/models/qwen3-how-to-run-and-fine-tun...