For some tasks, sure. But not for all tasks. And for some tasks, cost per token is irrelevant if it provides real benefits that are oom compared to what you had.
Local models are indeed becoming "good enough" for some tasks, but there are still tasks that they can't touch. There's a recent benchmark for kernel writing. Fable wrote a kernel that provides ~30% more throughput per unit of compute compared to the latest Opus max / gpt max. Does it matter how much that session cost in terms of one session if you can take that kernel, deploy it on your inference fleet and "magically" get 30% more tokens served to your clients? There are companies that would pay millions for such a "leap". Because they can make more millions down the line.
How many software developers were working on code like the one you describe?
The problem is going to become that there's no incentive for anyone to run the stupidly-expensive training phase. May God have mercy on the stock market.
That's so obviously not true that I don't even think it's worth the energy to even debate it. It's been said for years, yet here we are, constantly improving. People really don't get RL / the bitter lesson, do they?
> It follows that in a few generations, open models inferencing will be about as good as closed model inferencing.
Not a chance. There's hundreds of billions of dollars on one side, and oom less on the other. There's also scaling laws and information theory. No matter how good, a 30B model will not be able to be better than a 3T+ model, all things being equal.
You are mistaking models becoming "good enough" for an increasingly number of tasks, which I agree is happening, with SotA models stagnating, hitting walls etc. That will not happen for many many years to come.
Curious where you draw this conclusion from? Most benchmarks still show continual steady progress https://metr.org/time-horizons/
I feel like the market is just gonna become way more mixed. Not like "everyone is switching to open" or "everyone is staying on frontier" but a mix of multiple models for different scenarios