I think the answer lies in the "we actually care a lot about that 1% (which is actually a lot more than 1%)".
Purely according to Artificial Analysis, Kimi K2.5 is rather competitive in regard to pure output quality, agentic evals are also close to or beating US made frontier models and, lest we forget, the model is far more affordable than said competitors, to a point where it is frankly silly that we are actually comparing them.
For what it's worth, of the models I have been able to test as of yet, when focusing purely on raw performance (meaning solely task adherence, output quality and agentic capabilities; so discounting price, speed, hosting flexibility), I have personally found the prior Kimi K2 Thinking model to be overall more usable and reliable than Gemini 3 Pro and Flash. Purely on output quality in very specific coding tasks, Opus 4.5 was in my testing leaps and bounds superior of both the Gemini models and K2 Thinking however, though task adherence was surprisingly less reliable than Haiku 4.5 or K2 Thinking.
Being many times more expensive and in some cases less reliably adhering to tasks, I really cannot say that Opus 4.5 is superior or Kimi K2 Thinking is inferior here. The latter is certainly better in my specific usage than any Gemini model and again, I haven't yet gone through this with K2.5. I try not to just presume from the outset that K2.5 is better than K2 Thinking, though even if K2.5 remains at the same level of quality and reliability, just with multi modal input, that'd make the model very competitive.
If we consider what typically happens with other technologies, we would expect open models to match others on general intelligence benchmarks in time. Sort of like how every brand of battery-powered drill you find at the store is very similar, despite being head and shoulders better than the best drill from 25 years ago.
Yes, as long as that gap stays consistent, there is no problem with building on ~9 months old tech from a business perspective. Heck, many companies are lagging behind tech advancements by decades and are doing fine.
They all get made in China, mostly all in the same facilities. Designs tend to converge under such conditions. Especially since design is not open loop - you talk to the supplier that will make your drill and the supplier might communicate how they already make drills for others.
But if it's just 33% as good, I wouldn't bother.
Top LLMs have passed a usability threshold in the past few months. I haven't had the feeling the open models (from any country) have passed it as well.
When they do, we'll have a realistic option of using the best and the most expensive vs the good and cheap. That will be great.
Maybe in 2026.
If you are a business there is no point trying to save $50 a month, as the cost of the engineers time is much greater.
It is highly dependent on what the best represents.
If you had a 100% chance of not breaking your arm on any given day, what kind of value would you place on it over a 99% chance on any given day. I would imagine it to be pretty high.
The top models are not perfect, so they don't really represent 100% of anything on any scale.
If the best you could do is have a 99% chance of not breaking your arm on any given day, then perhaps you might be more stoic about something that is 99% of 99% Which is close enough to 98% that you are 'only' going to double the number of broken arms you get in a year.
I suspect using AI in production will be a calculated more as likelihood of pain than increased widgets per hour. Recovery from disaster can eat any productivity gains easily.
Nvidia was noted as having most of their revenue heavily tilted against only a few major customers for example.
i.e. whether you go for "only inference" or "both" it'll work out similar somewhere.
Kind of why Nvidia see themselves forced to make such deals, to lock labs in before they loose another to objectively superior inference.
Staying purely with US labs, they just lost Meta and both Anthropic and Gemini models have been available on Vertex without relying on GPUs for a while. OpenAI equally is turning towards Cerebras so yeah, those deals will mean little for inference.
Those circular deals aren’t a thing for the inference providers I tend to go for and they don’t change the fact that using a full GPU purely for inference is like using a pickup truck for a track day. It can be done and the flexibility has advantages, but once you focus on one task, there are better tools for the job.
To be clear: The models are open weights everyone can simply download, because many labs publish them as such. The providers in question are generic hosters. It's the same as if you would get some managed wordpress hosting somewhere.
Start with where the GPU and rest of it comes from.
I personally have been enjoying shadeform to build the GPU setup I like.
Their cost is not real.
Plus you have things like MCP or agents that are mostly being spearheaded by companies like Anthropic. So if it is "the future" and you believe in it, then you should pay a premium to spearhead it.
You want to bet on the first Boeing not the cheapest copy of a Wright brother plane.
(Full disclosure, I dont think its the future and I think we are over leveraging on AI to a degree that is, no pun intended, misanthropic)
So what ?
The long view is to see the microcontroller as a commodity piece of hardware that is rapidly changing. Now is not the time to go all in on betamax and take 10 years leases on physical blockbuster stores when streaming is 2 weeks away.
Ai is possibly the most open technological advance I have experienced - there is no excuse, this time, for skilled operators to be stuck for decades with AWS or some other propriety blend of vendor lock-in.
I'll also note that there's zero vendor lock-in in either scenario. It's a simple question about the tradeoffs of indirect parasitism within the market. I'm not even taking a side on it. I don't even know for certain that any given open weights Chinese model was trained against US frontier models. Some people on HN have made accusations but I haven't seen anything particularly credible to back it up.
[1]: https://finance.yahoo.com/news/nvidia-accused-trying-cut-dea...
[2]: https://arstechnica.com/tech-policy/2025/12/openai-desperate...
[3]: https://www.businessinsider.com/anthropic-cut-pirated-millio...
So I suppose the AI companies employ all those data scientists and low level performance engineers to what, manage their website perhaps?
It's poor form to go around inserting your pet issue where it isn't relevant.
So really, the argument pretty well makes itself in favour of the $0.5 micro controller.
there are pretty good indications that the american llms have been trained on top of stolen data
This works with every novel I've tried so far in Gemini 3.
My actual prompt was a bit more convoluted than this (involving translation) so you may need to experiment a bit.
They can’t even officially account for any nvidia gpus they managed to buy outside the official channels.
How do you even do that? You can train on glorified chat logs from an expensive model, but that's hardly the same thing. "Model extraction" is ludicrously inefficient.
I am not going to comment on how they did it. But they were openly accused by OpenAI of it. I believe the discussion is over destillation vs foundational models.
https://www.jdsupra.com/legalnews/openai-accuses-deepseek-of...
There are other theories like OpenAI inflated their training costs to seek further investment in later growth quarters. Meanwhile Deepseek under reported their cost to portray China as more cost efficient investment. If that was the case then their performance is similar, with similar training costs but one side reported even the coffee from the coffee machine in the office in the total while the other only counted the minimal CPU cycle cost and not the GPU, energy, engineering etc. Which is plausible too.
I have no dog in the fight but the first accusation seemed quite serious, hence why I asked