Why? It's much more efficient to have centralized special purpose hardware to run enormous models and then ship the comparatively small result over the internet.
By analogy, you don't have a search engine running on your phone right?
Why? It's much more efficient to have centralized special purpose hardware to run enormous models and then ship the comparatively small result over the internet.
By analogy, you don't have a search engine running on your phone right?
Privacy, security, latency, offline availability, access to local data and services running on the device, just to name a few.
But in a few years we might be able to have LLMs running on our phones that work just as well if not better. Of couse as you mention the LLMs running on large servers might still be much more powerfull, but the local ones might be powerfull enough.
Will not happen any time soon. Consumer hardware can't even run GPT-4 locally, and won't be able for a looong time. Each GPT-4 instance runs on 8 A100. The cost of such system is ~$81K. Not even in the ballpark of what most consumers can afford.