If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.
If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.
It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.
Therefore, I believe we are nowhere near 'good enough'.
I never drive my steam engine to work these days. It isn't good enough.
To borrow your steam engine analogy, if local LLMs get as good as a Toyota Prius, even if OpenAI / Anthropic offer Ferraris, most people will be happy with their Prius as their daily driver.
Similarly, if the big labs start raising prices or cutting usage, you won't be able to use it as much as you want -- whereas a local LLM will run all day every day without costing you any extra money.
So right now you are right, but who knows how long that will last.
This is not nearly true for everyone else in the world.
For example, think about the world in ~2021 pre-LLM. Would anyone say the sentence "I only want the fastest and smartest humans working on my project"?
No of course not. Most people don't want to pay $10 million dollar salary to the best programmers in the world. They prefer to pay $200k salary to a median programmer and that's good enough for their ecommerce website.
But so many teams said they wanted to Raise the Bar to infinity and hire a World Class Team.
I can see my new Thursday afternoon "oh chit" moment being that I didn't complete my weekly task because i torched all of those tokens M-W doing task/ticket grooming using the hot hot model instead of the dodo model with jira mcp connector :D
However, there are many use cases where they aren’t the right tool for the job.
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
I don’t expect the economics of local vs cloud ai to change until either the bubble pops or new ai chips land that can run big models fast with low power demand.
You just can’t compress 1.5T into 4B without losing useful stuff.
But that doesn’t matter that much. If 4B can compose tools, retrieve info, and not get stuck in loops that is going to handle a lot of use cases
Think you missed a word there.