Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
Edit: removed a comment that was uncharitable and rude, for which I apologize.
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
We've gone from 80% in some places to 80% in some more places.
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
Have you used recent ones on any large projects or coding issues? They've improved tremendously lately in my experience, in terms of implementing non-trivial things, debugging, handling legacy/complex codebases, etc. I usually keep a text document of things that models failed to do/fix and recent ones have wiped it all away.
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
People were talking about plateau for years already.
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
never mind that theres no guarantee we'll get that mythical AI. Never mind that the societal reformations would also impact their revenue numbers.
do you think it will be exponential forever?
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)
What a time to be alive.
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
Replacing white collar work would be worth dozens of trillions of dollars at minimum. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B Annual Revenue Run Rate in 200 days. And Anthropic is growing so fast they're on pace to hit $100B in Annual Revenue.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic and OpenAI are still growing enterprise usage heavily. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit about what qwen does. And specialized models often perform worse than generalized ones.
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
What distinction are you drawing?
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
It's not even anything controversial..
It’s not happening and the imminent bust is coming. Strap in while the music gets turned up (dots, ipo etc) and people decide to leave the partayyy!
Meanwhile, in the US, companies like Boeing, GM, and Intel will never be allowed to experience more than minor financial inconvenience before the government bails them out with protectionism, guaranteed loans, and outright subsidies.
I don't see a material difference between how the Chinese government treats their strategically-important industries and the way we do here in the West. Terms like "socialism," "Communism," and "capitalism" are just fodder for Fox News camp-followers.
it's also why there have been so many calls for regulation and slowdowns.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
I remember when bandwidth was super expensive and now it’s dirt cheap.
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.