And before that, businesses will be able to get decent results with dedicated inference hardware.
I could absolutely see local AI taking the place of voicemail/call screening completely. Call me and you get my AI, who will route the call to me if approved, or even choose to service the call itself. If a friend of mine calls who doesn't have $20K to spend on a great AI rig, they would certainly have permission to steal otherwise wasted cycles from mine. An answering machine isn't much different than an issue tracker, and some of those are 97% AIs having perfectly intelligible conversations with each other, and 3% humans being eagerly serviced by AI.
The old objection was "this is so hard to set up, nobody wants to run a server!" Now it can set itself up.
tldr; you don't have to rent from some data center. You can get high utilization out of a local box. This 1) puts a hard ceiling on what a data center can charge, and 2) they won't be able to compete on privacy, which locally can be complete.
The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
Agentic AI has pluses and minuses for cloud efficiency. The plus is that usage could be very bursty, but the minus is that agents will more fully utilize a local system. The main disadvantage of local AI is that you would be paying a large amount for a system mostly doing nothing. If it's constantly working on different projects and integrating that data, you get use out of every penny that you spent. Every GPU you added would instantly make the thing smarter.
What's more, your local AI could offload an agent to the cloud if it needed to. It could do this rationally, based on your personal desire for privacy.
The abstract benefits are quickly outstripped. The same way adding more highway lanes never improves gridlock. Its inducement.