Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.
> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)
And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.
That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.
Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.
You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.
Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.
Yes? Just like every single model from every single AI lab.
A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?
Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.
Wait a second, are we taking into account the massive difference in terms of resources of these two companies?
You're either competitive or not.
Yes.
Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.
Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.
It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.
Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.
Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.
Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.
So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.
The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.
The host's survival is not the benchmark of the reproductive unit, which are the cancer's cells short-term reproduction.
The distinction is the reason benchmaxxing is in fact not good: the benchmark is only approximately related to what people actually care about. You do want your cells to reproduce sucessfully, after all; you just also want some emergency stop buttons for when they go wrong, and those things failing is your body's benchmark rather than your cell's benchmark.
(There's at least two examples of cancers that can be spread from host to host; lupine genital and taxmanian devil nasal, IIRC)
The truth is much more mundane. It's just Goodhart's Law.
Some might say their job is optimizing ways to make numbers appear better without substantial change in input or output.
These firms are literally hiring professionals from all fields to teach procedure
To teach processes that can subsequently be done agentically or in automated chains
Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth
So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?
They are all gaming these benchmarks, it is perfectly reasonable not to trust any of them.
Probably
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
Also a lot of questions to benchmark because opus 5 is completely useless model right now.
I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
Seriously.
First, almost all models are within spitting distances of eachother.
Second, it never translates to being better for my own workloads.
You just need to make your own benchmarks.
An aside: When did talking like incels became cool?
While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.
Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.
Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.
I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!
Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.