Data at https://gertlabs.com/rankings
613 karma · joined December 22, 2025
Auto-scaling RL, AI evaluations.
Data at https://gertlabs.com/rankings
This has the least measured skill differentiation of all of our environments, and not because forecasting/markets don't require skill or intelligence. Even the best models are so far from anticipating the behavior of the other agents and understanding the emergent effects that a 2025 model with a naive strategy can often outperform over the timeframes of the simulation simply because some other models in the simulation chose a similar self-reinforcing strategy. This likely happens to some degree in real markets.
You can watch these simulations here https://gertlabs.com/spectate?game=market
- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).
- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost.
- Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department.
The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series.
Data at https://gertlabs.com/rankings
We tested Mercury 2.5 Preview, which is nowhere close to the frontier (and not advertised as such), but it's actually usable as a general-purpose chatbot. It's comparable in problem solving ability to some last-gen open weights models, and the price and cost make it compelling. However, they have not figured out general purpose tool use and agentic coding (their model performs worse on our problems when given a custom harness). If they do, I see a lot of real-time applications that the speed and cost will enable.
But you don't need any kind of insider information to see how fast the world is changing. ChatGPT launched less than 4 years ago and the advances in robotics, unsolved maths, and software are all riding the steepest exponential improvement curve any of us have seen. Interesting times we live in.
I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.
We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks.
It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1].
Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI.
Data at https://gertlabs.com/rankings
(1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
Data at https://gertlabs.com/rankings
This trend has been there since we started evaluating models using different languages in February 2026 and if anything, the disparity has grown in frontier models. Even Google models prefer Kotlin/C#/Rust for coming up with creative ideas (compilation success is a different story). Data at https://gertlabs.com/rankings
That being said, models love to recommend Go, and Go does have a lot going for it, especially if you are serving a public-facing website. So most of our public facing API handlers are written in Go, and we offload some of our most important binaries to Rust. There are just too many reasons not to use the languages that models think a little more effectively in.
So just adding a language or tag filter can result in some pretty small sample sizes. You can see how many samples survived in the box plot, but that's probably bad UX that most people never see. There's a reason no other benchmark provides this type of data (even for our sample sizes it runs almost 10K USD/month to keep up to date with new releases).
Might be a good idea to reduce the ability to apply filters into a cohort with less than ~20 samples -- not the first time we've gotten that feedback. Seems like adding too many options to see individual sample variation is just misdirecting. I'm a nerd who loves data so I hate removing access, especially since the aggregate performance is very interesting (averaged across all languages, we see consistent and interesting performance data across models, like models outperforming with strongly typed languages). But tbh I think you're right and we'll try limiting filters to where we actually have statistically significant data.
I agree that Opus 5 is not a great model, despite being clearly intelligent. It seems like a personality problem in user-driven agentic coding workflows, not a real capability issue. Not incorporating unspoken user intent, going off topic, incorporating some of the pedantry you find in GPT 5.x models, etc.
That's also likely why Opus 5 ranks low on our "Social Intelligence" benchmark (https://gertlabs.com/rankings?mode=decision), although sample sizes on this one are still low.
Data at https://gertlabs.com/rankings
All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.
Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)
Data at https://gertlabs.com/rankings
We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.
Data at https://gertlabs.com/rankings
Agentic coding is what's most relevant to software engineers, but so is speed, which is a real usability issue right now. Third-party inference providers like Fireworks have bridged the gap for some previous releases.
It's exciting to see the open weights frontier becoming the norm. Competing on having the frontier model is going to be an increasingly difficult business.
I have not broken down the comment/code ratio, but that's actually a really interesting idea for a metric.
I would also like to test Cursor, but our policy is to only test models available on public routers for now.
Agentic coding data: https://gertlabs.com/rankings?mode=agentic_coding
However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore.
Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol.
Fable scores slightly lower, but with an average solution size of 12.2 KB.
Sonnet 5's performance is comparable to GLM 5.2 in both one-shot coding and agentic ability. However, it's about ~20% less verbose than GLM 5.2 in average code submission sizes, and uses fewer reasoning tokens, which reduces the cost gap and suggests it writes cleaner code. In practice, Sonnet 5 ends up being 40% more expensive and ~2x faster than GLM 5.2 in our evaluations (not 300% more expensive as the per-token pricing would suggest). Granted, GLM 5.2 is an extremely reasoning heavy model.
Overall, it's a solid release that gives Anthropic some standing in the price-conscious inference market.
Data at https://gertlabs.com/rankings
It doesn't have a higher capability score than Fable, though. We break our coding evaluations into 2 parts, and "one-shot coding" makes up part of the index, where Fable significantly outperforms every other model, which is why it's ranked at the top despite Sonnet 4.6 having a slightly higher median (and lower average) in long-horizon agentic workloads. One-shot coding tends to be the most correlated with other companies' model cards, whereas agentic coding is partly about how well a model can adapt to a custom harness. Fable also refused some tasks.
Data at https://gertlabs.com/rankings?ow=1&mode=oneshot_coding
We now use Sonnet 4.6 for a number of internal use cases we wouldn't have considered otherwise.
We test 11 popular/interesting languages (you can see the Languages chart to filter), but not Elixir -- although other evaluations have found that many LLMs solve more problems when working with Elixir [0]. Why models write code well in some languages over others seems to go beyond pre-training data (Python scores quite low for most models) and we don't fully understand it.
[0] https://elixirforum.com/t/llm-coding-benchmark-by-language/7...