This is an amazingly ignorant thing to say given the current pace of progress.
This is an amazingly ignorant thing to say given the current pace of progress.
And there are benchmarks that cleanly separate the SOTA models:
Saturation of benchmarks is a property of benchmarks just as much as of the models.
- one is scaling laws, where we found years ago that pretraining validation loss scales in an almost miraculously predictable way with data volume and compute. There are apparently theoretical bases for this that I don’t quite understand but this property alone is holding at every scale we’ve ever tested. There is not just “one” scaling law but the point is there are scaling laws and they continue to faithfully predict the performance gains we see
- one is benchmarks, which I always point to epoch capability index as a good summary of them in aggregate which makes it nice to plot on one curve the capability improvement over time
To me either one without the other is substantially weaker, the fact that theory and empirical measurements give you a very good scaling law on a more unintuitive quantity (pretraining validation loss) that’s only indirectly related to the downstream performance you care about, benchmarks (in aggregate) are more direct measures of downstream performance but are harder to nail down clean and well motivated “laws” from theory (as far as I can tell). Nevertheless we do in fact see a clear trend that is not slowing.
That doesn’t mean there aren’t a whole host of benchmark problems that don’t impact the numbers involved here (leakage from training data, fundamental flaws in the design, benchmaxxing) but they don’t change the larger story. These problems don’t plausibly explain the clean trends we see.
Even if both aren't true, your evidence was people saying two opposing things. The truth (if there is a single objective truth on a given thing) has little bearing on whether or not different people agree on it.
Something being "non-deterministic" is orthogonal to whether or not skill plays a role.
I think this is due to rapidly rising expectations.
When LLMs first show they can do some new thing, we're excited at first. Then, we quickly start taking it for granted, and get upset whenever the LLM fails.
Just three years ago, LLMs could barely hold a conversation. Now, they're writing entire code bases and solving famous mathematical conjectures, but we still focus on whatever they can't do.
In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".
(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. That would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)
I find even Deepseek Flash v4 0731 even outperforms Opus for me (at 10x the speed two).
Using Sol, Fable, Kimi 3 and other recent models has been unbelievable for me. I didn’t think we’d get to this level for years.
I’m using them for Ruby, TypeScript and Python. In large existing codebases but also lots of tiny tools.
Point being: it’s overall a worse experience even if the model is technically better at a lot of things.
We are not even close to what AI slowdown looks like.
The whole business side of things are now building Agentic Layer for Business applications. All of this Agentic Layer needs to be build and its happening right now and still needs a little bit of time.
Anthropic and co have the biggest and centralized reinforcement loop on the planet: Millions of people telling them what is good and what not due to thumbs up/down.
And for sure when the businesses are building the agentic layer they might give direct feedback to them.
While in parallel LLMs get better, more generic and a LOT cheaper too.
I mean the token prices in general as certain services were never really using a subscription.
I do run a claude subscripton right now though and since there capacity change, i hit the limit rarely in comparision to the past, but I don't think this will stay as it is.
Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.
Feel free to be a dick to someone else.
We all learned more from the prior comment than we did from this one.