What is expensive is model development and training, but again, that's nothing inherent to technology. It's just that Anthropic and other western vendors made a business decision to use cheap investors money to brute force a new model development and market grab instead of investing in cost optimizations for model training and inference like Chinese model developers and providers.
The question is if the MoE advances and compression (TurboQuant) makes common computers fast enough to run local models that are "Good Enough".
If anything breaks on my pc, it's going to be really rough.
My kids grow fast,a few years is a lot for them too.
And i have a bunch of games I want to play with my wife that require 2 computers sadly
Realistically Apple needs to put a few teams on it, and have built in support for running LLMs on OSX.
I assume if you integrate this on an OS level you might be able to pull off some additional tricks. So far most of the tools for running LLMs locally have been driven by volunteers or very small startups.
Microsoft could also do this, but they're too busy trying to upsell everyone. Microsoft wants to sell you on paying for each token.
Microsoft has spent a lot of time telling people to buy AI ready PCs.
I just don't think it's in their business model to offer on device LLM. They want you to subscribe to Copilot. Which encapsulates the entire enshitication of Windows. Microsoft demands more money. Subscribe to gamepass, subscribe to OneDrive!
I can see Microsoft pushing the subscription angle, but I do think there is a place for on-device capabilities since those will always be more responsive for small tasks.
I believe a few small companies have already specialized in this. But imagine a situation where maybe you pay a small subscription and Open AI allows you to download non open source models guaranteed to work to a degree.
It's going to be an interesting decade.