The reason why local LLMs are unlikely to displace cloud LLMs is memory footprint, and search. The most capable models require hundreds of GB of memory, impractical for consumer devices.
I run Qwen 3 2507 locally using llama-cpp, it's not a bad model, but I still use cloud models more, mainly due to them having good search RAG. There are local tools for this, but they don't work as well, this might continue to improve, but I don't think it's going to get better than the API integrations with google/bing that cloud models use.
They also already practice a concept of computational offloading with the Apple Watch and iPhone; more complicated fitness calculations, like VO2Max, rely on watch-collected data, but evidence suggests they’re calculated on the phone (new VO2Max algorithms are implemented when you update iOS, not watchOS)
So yeah; I can imagine a future where Apple devices could offload substantial AI requests to other devices on your Apple account, to optimize for both power consumption (plugged in versus battery) and speed (if you have a more powerful Mac versus your iPhone). There’s good precedent in the Apple ecosystem for this. Then, of course, the highest tier of requests are processed in their private cloud.
And don't worry, I'm sure some enterprising electricity company is working out how to give you free electricity in exchange for beaming more ads into your home.
* If you trust the OS vendor, why wouldn't you trust them to handle AI queries in a responsible, privacy respecting manner?
* If you don't trust your OS vendor, you have a bigger problem than just privacy. Stop using it.
What makes people think that on-device processed queries can't be logged and sent off for analysis anyway?
I envy your very simple, sedentary life where you are never outside of a high-speed wifi bubble.
Look at almost every Apple ad: It's people climbing rocks, surfing, skiing, enjoying majestic vistas, and all those things that very often come with reduced or zero connectivity.
Apple isn't trying to reach couch potatoes.