The only thing preventing this switch from starting in earnest is the data center buildout monopolizing all current and future GPUs
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
You just can’t compress 1.5T into 4B without losing useful stuff.
But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases