There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
There's no way large companies outside the US will pay the "US AI lab" premium if they can get the same workloads done at a fraction of the cost using open-weight models that they can self-host and optimize/fine-tune on.
For simple queries, we have reached the threshold since the beginning of the year, and models are good enough from every provider to make a meaningful difference between one another. (ChatGPT, Claude, Gemini, Grok, MuseSpark, Kimi, DeepSeek, GLM...)
The real unlock will be, and you can already see it with GPT-5.6 and Fable-5, to delegate complex enough tasks that will take more than 24 hours to get done and they will not lose track. I'm not talking about a loop, but the actual intelligence to recover from these compounding errors that accumulate in dumber models.
We're still a long way from the intelligence needed to let one of these agents go ahead and supervise multiple layers of sub-agents underneath to do complex orchestration. The future looks very promising and exciting. Imagine having the possibility of a Frontier model orchestrating as many sub-agents as needed that are running on cheaper models like DeepSeek.
While SOTAs handle these errors better, they compound in all models and there's a term for that. It starts with cluster and ends with an expletive.
I wish I could, but I don't see the need for human steering going away soon if the task involves anything novel (see Terry Tao's chat).
Yes it's much easier to have a smarter model that goes straight to the correct answer first, but it may not be necessary or economical. There's a minimum bar for the model where it understands problems and knows the right step to correct them, and above that newer models give diminishing returns.
That's basically ASI not AGI, if you agree humans are NGI (natural general intelligence) and make mistakes and wrong decisions in solutions all the time. Right steps with some wrong ones is acceptable though for AGI.
For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.
You use Fable 5 right? If that’s good enough for you now, why wouldn’t a Chinese model that’s as good as Fable 5 but at 10% the cost be good enough in 6 months?
I use Claude Code semi-heavily for my small business, and the $100/mo I pay for that is a rounding error compared to the value it provides.
If I can avoid spending an hour or two "massaging" the output from a lower-end model once, or it avoids introducing one load-bearing (sorry, couldn't resist) bug, then that's the entire $100 right there.
Hell, you could argue that the best "coding model" that we have at the moment is the human brain, and people will gladly pay $10,000/mo for one of them.
Arguing over $20 vs $100 for something that actually puts in work just seems insane to me.
Fable 5 is still going to mess things up at any sufficient complexity. The advantage of low cost models with "good enough" intelligence is they can recursively correct. Why? Because it is cheap. Proper requirements and tests and subagents take away increasing amounts of work, at a cost that is not prohibitive.
If you are reviewing code manually you might consider Fable 5 a worse option. As it articulates itself with higher confidence and you already know it is capable, you are may be more likely to miss a mistake. You know to be on guard with a junior engineer. Reviewing a senior who suddenly makes some weird stochastic mistake can be a lot harder. It would be like if the smartest human engineer you knew was capable of some random brainfart in the middle of their massive diff. Imo, much harder to deal with.
Of course, we should keep in mind Fable 5 is only expensive today. It will be cheaper in the future. Autonomous, recursive prompting and improvement is the clear end state. Especially for entities that will always have the budget for that at the SOTA frontier.
Which was an argument for using every less powerful model since the moment they got useful, right?
When was that? Opus 4.5 maybe? Let's say Opus 4.5 for the sake of the argument. So back then we were like "DeepSeek is not good enough, I need Opus 4.5". Now DeepSeek is better than Opus 4.5. So if Opus 4.5 was good enough back then, DeepSeek is better than that now.
Sure, it's always nicer to have a slightly better model. But the price difference starts mattering a lot more when all the models are already sufficiently good.
We've just spun up our first Hermes agent, with direct API access to our main inventory system and that's expected to find another few grand per month in misallocation/inefficiency.
I wouldn't be surprised if we were doing more like $10k/mo higher in 6-9 months' time.
When you're talking about numbers like this, the fact that one AI is $100/mo and another is $10/mo or $40/mo doesn't matter. They could make GLM-5.2, or any other Opus 4.5-class model free and it still wouldn't make sense to deploy in a commercial context.
The other angle I'd approach things from is that Opus 4.5 (and I'd agree with you that that model was the saddle point) was "good enough" for the types of things we were asking it to do back then, but as the models have become more capable the tasks we're asking them to do have also expanded with it.
I know I've personally gone from "hey can fix this race condition with a Redis mutex" 6 months ago to "Independently redesign this full embedded USB stack and QA it end-to-end, working around a specific Kernel bug in macOS Tahoe that requires decompilation to find the source of, while keeping in mind the constraints of our 8-bit AVR chip from 2011" now.
But that said, yes, maybe in 5 years' time we will reach an "intelligence saturation" where the average person won't be able to even conceive of how to use the new SOTA.
Surely they are using read-only access.
This is such a ridiculous objection for how beloved it is. Wide swathes of the public can't cope with any adversity or risk.
With an agent (especially incompetently employed), the danger of unwittingly destroying your company (or at least, the crucial data/reputation) is rather higher. We are notoriously bad at estimating the downside risks in complex systems.
The most obvious case is the downside risks in complex financial constructs... things look great for a while ... until a sudden surprising collapse arrives and totally destroys all the upside you think you have created.
In any case, GP's post was not such a balanced consideration; it was just parroting a beloved risk-aversion meme that can easily be deployed against building anything (what if the building falls on top of someone?) or even leaving home to go to work ("travelling in a hunk of steel at lethal speeds – let me assure you that absolutely nothing can go wrong here, mate.")
What I find tiresome about that meme is the presumption that "something can go wrong" is useful input on its own. It's not. Mistakes are made all the time, the only way to avoid that is to stop breathing. Even in the process of me standing up and going to the loo, something can go wrong.
If the guy wants to make a case that it's too dangerous for the expected benefits, he has to actually make that case. Saying "risk exists" with no elaboration is a waste of HTML. "something can go wrong" every time he swallows food, yet mysteriously he still does it.
(the suicide analogies may seem mean-spirited, but I kind of mean it. If you consider every action primarily from a standpoint of "what harm or irreversible change can result from this", the only permissible path is to do nothing. To be moral is to be as close as possible to a rock or another inanimate object.)
Are you saying we shouldn't care about the future of affordability and access because at this moment we have seemingly endless access?
Sounds extremely short sighted.
If $100 Claud Max subscription works for you, then great.
But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.
For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.
And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.
This is made brutally obvious by anthropics customer support for people with such accounts.
Right now the US dominates everyone else in actual chips in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.
It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.
That one is dirt cheap at API pricing, I can't imagine quota is going to be a concern on the $200 subscription, which in my opinion easily supports full time use of 5.6 Sol on xhigh.
A couple of very talented friends were uttering curses upon the entire bloodline of whoever convinced them to try letting sol xhigh do serious work. Deepseek cleaned it up for a fraction of $20.
The cost per task was $0.03 with DeepSeek, $0.05 with Luna. $1.23 for Sol.
Tokens per second 132, 202, 70 respectively.
They also previously said prices will go down significantly once they get a hold of the upcoming Huawei chips (later this year).
Prices are going up just because they can. It can easily come back down. They aren't strained by some IPO / VCs requiring them to 1000x their earnings.
Think about it this way.
Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time. To the lay person this sounds trivial but to a serious business this intelligence gap could represent millions, or billions of dollars.
But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.
There is a reason why Chinese open weight models are now popular even in American enterprises, because CTOs realize that they are indeed good enough.
The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.
And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder.
For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, probably until we have AGI that can replace employees, maybe even a bit after.
This is why I find the "good enough" arguments silly. Like, the usefulness of an AI tops out to you when you can pair program with it? Seriously? You cannot envision ways in which more advanced AI enables you to do more, better? That's bizarre to me. I don't ever see myself running out of problems to solve.
Is this what you are looking forward to?
This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.
On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.
Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.
Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.
My overall thoughts (released over some time):
https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...
https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...
https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.
That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.
> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,
GLM 5.3 bridges this gap.
> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.
Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!
> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.
Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're a lot larger e.g. Fable. It is insanely bad. Do you know how much more resources "Western" companies have? Most in China don't have random GPUs to "play with" like every "frontier lab" employee does.
> As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're born lucky though. Efficiency is "free". The next generation hardware e.g. Nvidia claims Blackwell -> Rubin is 10x efficiency (verified by Neoclouds apparently).
I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough.
I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have.
And I don't see this ending. When better models come out that take that from 80% to 99.x%, I will have that better model manage teams of other models and move up the abstraction layer.
If models get even better than that, perhaps I stop reviewing PRs entirely. Maybe normies can start using agents to build real things.
Unless your business doesn't have many problems to solve and isn't in a competitive environment, it will benefit from using the best models.
My point is you don’t need the best model if you just put in QA processes that can be done by models also. And if you don’t have that, the model is probably not going to be good enough.
Fable has only been out for a month but somehow everyone is supposed to have moved to a completely different way of working that supposedly only works for Fable and nothing else…
This kind of takes makes me cringe. Why don't you go back to LinkedIn?
I have no idea what you mean by “pair program with an agent”, but Opus have been able of autonomous coding since last November, and with any half-decent harness even local Qwen3.5 was able to do so 6 months ago.
Fable is a stronger model, which means it can solve harder tasks but it's also over-hyped, because only a small fraction of task is hard enough to be Fable-worthy.
Fable is the only one that reliably one shots complex changes and makes the right design choices. Everything else requires handholding.
I can let Fable loose on a 12+ hour (for AI) task and it will have performed it flawlessly when I come back the next day. K3 and Opus are not like this.
And no, our harness is not the limiting factor here.
This is not appropriate for HN. Please review the guidelines: https://news.ycombinator.com/newsguidelines.html
There are businesses other than FAANG. I don't know why this is so hard for FAANG employees to understand.
It's been 2 months since Fable was released to the general, man. Nobody knows what's going on inside of these companies except the people at the coal face.
It's funny to see that Anthopic shills have been saying the exact same thing for the past two years now (and it was OpenAI fans before). It's amazing to see that Claude 3 Sonnet was "great" but now that even Qwen 9B is better than this version of Sonnet DeepSeek V4 is still not good enough despite being stronger than Opus 4.7 was.
> Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time
If you think Fable makes 20 times fewer mistakes than DS4 you're delusional. It doesn't even do 20 fewer mistake than Gemma 4…
Same reason it makes sense to assign a team of humans that cost $100k/mo to a product that brings in $5M/mo, rather than one human with 5 Claude Max subs.
The cost is a rounding error.
This is why I quite like Kimi K3 - close to the same performance (definitely like Opus, approaching Fable), noticeably cheaper, generally good enough for me to daily drive. Only problem is that their official provider (on the Vivace plan) feels kinda slow, I'd say close to 2x slower than Opus on Max reasoning on average (probably more relatable than Fable).
You can say this about literally every product we buy. And yet...
So, death sentence even to frontier models?
Another one I did was a printer data stream translator from an obscure format to PostScript/PDF (or just PNGs), complete with cups support, etc so these old apps can easily be hooked up.
Flash is capable now of running long range defined-goal tasks like this.
At this point, I don't even know if its possible
- You describe the breakdown in terms of time but it's more accurately a function of reasoning complexity.
- You seem to assume that no intermediate evaluation is possible.
- Often it is (e.g. the build breaks or tests start failing), allowing for course correction. There's definitely a cost to that but it can still be cost effective if the accuracy is "good enough" and the price difference significant.
- There are numerous tasks that don't require Fable or GPT5.6 level reasoning to improve efficiency by an order of magnitude.
Such as Decision Making. /s
You just can't set a high enough threshold of intellectual effort for critical decisions.
Many devs who have never tried from either side build it all up in their head but it's almost always been a matter of thorough tedium, which LLMs are excellent at churning through, especially when there's api docs/headers/code comments.
If you have the space, try mirroring your port at the switch level and capturing every packet then making it go through them all to look for whatever. We have NSA at home lol
I've been working with DeepSeek V4 Flash 0731. I'd say that it's maybe not quite as smart as Opus 4.5, but it's willing to think things through carefully and keep going until it gets a good answer. So it's a decent Opus 4.5 replacement. Just let it cook.
It isn't Opus 5 or Fable 5. But it's nearly free on Open Router, and it's self hostable on a Mac Studio with plenty of RAM, or using an RTX Pro 6000 Blackwell or two. Which is chump change for any company that employs programmers.
It would absolutely have been a frontier model last December.
In what kind of sad and failed dystopia is this a "saving grace"? For whom?
For Anthropic and OpenAI, presumably. And the rather large economic distortion field around them, that may or may not go very badly for all our retirement funds if those firms become insolvent...
Well... I would think that the whole AI industry in the US are working towards public bailouts... Which I guess they'll get under the current administration... So they'll be fine... Nobody there really seems interested in actually creating a profitable business anyway...
I dont see how the outlook is any better for the open weight companies. They’re in the exact same situation as the closed weight companies except they have had much less revenue, and built up less of a brand, leading up the the point where they are equal in terms of model quality.
I joined a company that is an Anthropic shop and I am genuinely shocked.
Sonnet 5 is a little better on long horizon tasks and headless unsupervised agent workflows - but for in-IDE workflows, it's virtually unusable.
I am so used to flipping around my codebase at warp speed with DeepSeek flash. It's so fast and accurate I don't have time for parallel agents. It's a really rewarding workflow.
Moving to Sonnet, you ask is something simple like "split this into a seperate file" "implement this method" "this is my schema, implement a repository for it". It'll spend 30 minutes thinking and charge like $12. And no token caching, what are you even doing Anthropic?
It's unusable.
DeepSeek are in a league of their own
Like Europe?
First of all, no one knows the "true cost" of any of this, yet, but we know it's expensive. To what extent are the Chinese labs being subsidized? Are they real businesses?
Second, the Chinese labs aren't some "super geniuses", while the American labs are full of clowns. As of today, like the past 3 years, American labs are SOTA. That might change, but let's not act like the American labs don't know what they are doing.
The idea that people are going to use cheaper models for cheaper work isn't some novel revelation, it's completely obvious. People are doing it already, eschewing Fable.
The point is it takes money to keep developing models. Everyone is playing by the same rules. At this point, the US labs are trying to build businesses. I'm not really sure what the Chinese labs goals are. But I do know they aren't doing charity work.
And who gives a flying fuck. I am a "real" business and I count my money. It is not my life goal to prop some fat cats crying crocodile tears. Granted I do not use Chinese models. I use Junie straight from my JetBrain's IDEs that in turn uses Gemini Flash. Very cheap and more than enough for my use.
this has already happened with manufacturing, so it isn't surprising that other industries follow.
The US premium in engineering and scientific endeavors have been lacking for the past 30-40 years, and if it werent for tech and silicon valley, the US would have nothing state of the art. Even on that front, the US is falling behind given how much effort in tech has been diverted into privacy invading, and advertising.
The US has been riding momentum, but eventually that momentum will stop. It will take half a century to get back up to speed, and by then, the US will have fallen behind so far that catching back up will seem impossible.
The telling evidence would be if china has the first moonbase before the US does. I think this is highly likely looking at today's US administration.
I downgraded my Claude subscription and delegated my Claude Opus access to serve the role of an Architect to brainstorm and plan every step of development.
I leave the development to Deepseek.
Claude gets to review at many layers. It is often just as good as if I let Opus develop it by itself(the architect session will find similar number/level of gaps).
Deepseek flash as an architect and problem solver is not as thorough as Opus5 + high. Codex sol+ high is even better than Opus 5 at this moment for my needs.
Spirit was broken by oil prices which everyone pays the same for. (There is no cheaper jet fuel alternative).
Not a good comparison to the point of wrong conclusions.
Spirit was broken by oil prices because they target the low end customer with their ticket prices. Oil prices went up and they had no pricing headroom to charge more on tickets so they simply went kaput. Ryanair also suffered a fair bit. Other airlines did (comparatively) fine because they had the ability to increase prices since their customers are less price sensitive.
It’s a classical business lesson that being a “cost-sensitive” vs a “value-sensitive” business (what this tradeoff is called) is a tradeoff. It’s notably recommended that startups don’t target the lower end in prices since you can’t compete on cost with a business that has more economy of scale than you; you have to compete on features. And being a cost-sensitive business means that you are affected much more than other businesses by changes in material/component prices, because a 13-cent increase in the cost of a component matters more the more product you sell, and if you increase prices too much customers will start to wonder if the “budget” brand is really a good value proposition over the mid-end or high-end ones anymore.
Literally everything you're saying doesn't line up with reality. Apple's most popular product today are two lower end (neo and mini).
You make the mistake of thinking your theory defines reality, when in fact reality says quite the opposite here.
I'm saying this because your logic is the classic logic that everyting thinks makes sense until you do it. Everyone can't target the high end, there are limited customers with many choices, making it far harder actually to win.
You target the demand/pain/need regardless of market.
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
https://en.wikipedia.org/wiki/List_of_defunct_airlines_of_th...
Seems more like a overall industry problem, not limited to low cost carriers
At least here in Germany Aldi isn't even really limited to the poor, it's famously a place where you can run into anyone. Where I used to live in Berlin close to the government district I literally on occasion ran into the chancellor (and her bodyguards). Aspirational shopping where you buy premium goods to pretend to have higher social status honestly seems a bit on its way out. Even middle class people seem to consciously shop more utilitarian now.
The ByteDance folks are apparently training a mythos level model 10T params apparently. If they do would it still be subsidized at these cheap rates?
The bet is on using AI to gain competitive advantage. You don't win the stock market or make the deadliest drone by switching to the cheap model
Really? How many times a small team has outperformed a much bigger one just because they were "doing it right"?
I have been in software companies where most software produced was bad. Not just the code, the overall design everywhere. So... bad engineers with the most expensive model, or great engineers with cheaper models?
> You don't win the stock market or make the deadliest drone by switching to the cheap model
The question is not "can you win with the best model?", it is "can you not win without the best model?".
I have been using it since it got released and its as good for scoped coding tasks, as the other big models I use, but just soo much cheaper.
As it is now clear by behaviour of companies and US govt, all these investments will be backstopped by US govt. No US AI company will go hungry, they are national champions.