HNHacker News
TopNewBestAskShowJobs

gertlabs

613 karma · joined December 22, 2025

gertlabs.com

Auto-scaling RL, AI evaluations.

submissionscomments
gertlabs··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.

Data at https://gertlabs.com/rankings

gertlabs··on Artificial intelligence now beats some of the best human forecasters
We measure skill differentiation between frontier / last-gen LLMs across our environments, and one of our curated coding environments is a closed-system market simulator, containing only other agents and some system participants (a market maker and a liquidity provider via issuance / buybacks) whose behavior is fully defined for all of the agents.

This has the least measured skill differentiation of all of our environments, and not because forecasting/markets don't require skill or intelligence. Even the best models are so far from anticipating the behavior of the other agents and understanding the emergent effects that a 2025 model with a naive strategy can often outperform over the timeframes of the simulation simply because some other models in the simulation chose a similar self-reinforcing strategy. This likely happens to some degree in real markets.

You can watch these simulations here https://gertlabs.com/spectate?game=market

gertlabs··on DeepSeek v4.1 Flash Is Now Our Best Hacking Model
We ran v4.1 Flash through our evaluations and found it to be smarter and faster than V4 Flash, with a commensurate price bump. Some notes:

- Because it's a heavy reasoner, it sits near Gemini 3.7 Flash on the Pareto front (not as cheap as the price suggests in practice).

- Closer than expected to the top open weights models (GLM 5.3 and Kimi K3) in agentic coding, at lower cost.

- Chinese models have always been strong iterators in an agentic harness. This model is no different, reaching an average percentile ~20% higher when given a harness vs a one-shot solution. That one-shot fluid intelligence is what makes a model feel smart, though, and typically results in fewer attempts/tokens to solve a problem, and American frontier models are still far ahead in that department.

The new architecture is interesting. It puts pricing between their old Flash and Pro lineups, suggesting they might be abandoning their super-cheap flash models (which weren't that fast due to heavy reasoning) and their pro models (which sort of flopped and weren't consistently better than their flash models, despite the size/cost) and shipping a strong intermediate that competes with the Gemini Flash series.

Data at https://gertlabs.com/rankings

gertlabs··on Mercury 2.5
Inception is one of the most interesting neolabs with their diffusion-based architectures. My understanding is that their primary business is low latency voice applications but they are seriously pursuing coding.

We tested Mercury 2.5 Preview, which is nowhere close to the frontier (and not advertised as such), but it's actually usable as a general-purpose chatbot. It's comparable in problem solving ability to some last-gen open weights models, and the price and cost make it compelling. However, they have not figured out general purpose tool use and agentic coding (their model performs worse on our problems when given a custom harness). If they do, I see a lot of real-time applications that the speed and cost will enable.

gertlabs··on An Alien Mind
The results we've been seeing internally on our physics and circuit design environments are expert-level and beyond-expert-level results from models that Astra completely outclasses across the board on our evaluation suite (Fable 5+/Opus 5/Grok 4.6 were all worthy of being called AGI in my opinion). That's hard tech that will translate to real product innovation.

But you don't need any kind of insider information to see how fast the world is changing. ChatGPT launched less than 4 years ago and the advances in robotics, unsolved maths, and software are all riding the steepest exponential improvement curve any of us have seen. Interesting times we live in.

gertlabs··on An Alien Mind
> Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.

I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.

We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks.

It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1].

Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI.

Data at https://gertlabs.com/rankings

(1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)

gertlabs··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.

We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.

Data at https://gertlabs.com/rankings

gertlabs··on Ornith-1.5: From Self-Scaffolding to Self-Improvement
These models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
gertlabs··on Go is an ideal language for AI-assisted software engineering
We've seen a pretty consistent pattern in our evaluations where Go is among the languages that models perform worst with (alongside Python), for reasons unclear. Our coding evaluations are typically measuring the foresight and planning expressed in code that is run in interactive environments/games.

This trend has been there since we started evaluating models using different languages in February 2026 and if anything, the disparity has grown in frontier models. Even Google models prefer Kotlin/C#/Rust for coming up with creative ideas (compilation success is a different story). Data at https://gertlabs.com/rankings

That being said, models love to recommend Go, and Go does have a lot going for it, especially if you are serving a public-facing website. So most of our public facing API handlers are written in Go, and we offload some of our most important binaries to Rust. There are just too many reasons not to use the languages that models think a little more effectively in.

gertlabs··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
SlimeBallBench? Looks cool, love the implementation! You would most likely need a LOT of environments like these if your goal was selling them to labs
gertlabs··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
The reality is that cost is the primary constraint for the public benchmark we provide. While we run enough samples to get results that are generally quite accurate on average, we only produce ~10 coding submissions per language for each model and those are across random environments, which naturally has noise. Plus that's split between agentic coding sessions and one-shot coding.

So just adding a language or tag filter can result in some pretty small sample sizes. You can see how many samples survived in the box plot, but that's probably bad UX that most people never see. There's a reason no other benchmark provides this type of data (even for our sample sizes it runs almost 10K USD/month to keep up to date with new releases).

Might be a good idea to reduce the ability to apply filters into a cohort with less than ~20 samples -- not the first time we've gotten that feedback. Seems like adding too many options to see individual sample variation is just misdirecting. I'm a nerd who loves data so I hate removing access, especially since the aggregate performance is very interesting (averaged across all languages, we see consistent and interesting performance data across models, like models outperforming with strongly typed languages). But tbh I think you're right and we'll try limiting filters to where we actually have statistically significant data.

gertlabs··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
That's a different problem than benchmark saturation, and it's something that we are actively working on measuring objectively.

I agree that Opus 5 is not a great model, despite being clearly intelligent. It seems like a personality problem in user-driven agentic coding workflows, not a real capability issue. Not incorporating unspoken user intent, going off topic, incorporating some of the pedantry you find in GPT 5.x models, etc.

That's also likely why Opus 5 ranks low on our "Social Intelligence" benchmark (https://gertlabs.com/rankings?mode=decision), although sample sizes on this one are still low.

gertlabs··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better than any non-aggregator benchmark, and likely at lower cost to run.

Data at https://gertlabs.com/rankings

gertlabs··on ARC-AGI Leaderboard
We run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude.

All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.

Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)

Data at https://gertlabs.com/rankings

gertlabs··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
The Efficiency tab at https://gertlabs.com/rankings?mode=oneshot_coding (only have cost data for the coding evaluations)
gertlabs··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
Yes, they are all benchmaxxed, but the question is how benchmaxxed they are relative to each other.

We run an evaluation that only compares models in open-ended multi-agent environments where agents affect each other, primarily testing writing code. It's designed to be less vulnerable because there's no solution set, and it's been pretty effective and tends to rank Chinese models lower than their advertised model cards (relative to US models). Kimi K3 is a bit of an exception there -- it truly is a near-frontier model. But it's so slow. Ironically, Muse Spark 1.1 is one of the strongest models we've tested after Fable and Sol while also leading the cost efficiency curve. Big turnaround from Llama 4.

Data at https://gertlabs.com/rankings

gertlabs··on Moonshot AI suspends new subscriptions due to Kimi K3 demand
In our multi-agent game coding evaluations, we usually see Chinese models struggle in one-shot reasoning but make up for it with tool use and iterating towards better solutions. Kimi K3 follows that pattern, ranking 19th in one-shot coding and 3rd in agentic coding (where the model gets a harness and tools and many calls to iterate toward better code). Only Sol and Fable have better average agentic coding submissions.

Agentic coding is what's most relevant to software engineers, but so is speed, which is a real usability issue right now. Third-party inference providers like Fireworks have bridged the gap for some previous releases.

It's exciting to see the open weights frontier becoming the norm. Competing on having the frontier model is going to be an increasingly difficult business.

Data at https://gertlabs.com/rankings?mode=agentic_coding

gertlabs··on Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
LLM as judge / self-distillation is effective insofar as it can make models more reliably do things they are already capable of. But I agree that for pushing the frontier of what a model is capable of understanding and producing, incestuous is a good word and it's unlikely to scale far.
gertlabs··on GPT-5.6
The human solutions are all written in Python, which creates a significant length bias, whereas the AI models are assigned to create solutions randomly distributed across 11 relevant programming languages, most of which are inherently more verbose than Python.

I have not broken down the comment/code ratio, but that's actually a really interesting idea for a metric.

I would also like to test Cursor, but our policy is to only test models available on public routers for now.

gertlabs··on GPT-5.6
Gemini models struggle with agentic coding/tool use/exploration, but they are actually quite smart in one-shot reasoning. They're not as far behind as people think. It's mostly post-training and productization issues, which are easier to fix than pre-training/mid-training issues.

Agentic coding data: https://gertlabs.com/rankings?mode=agentic_coding

gertlabs··on GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
That's a much shorter and more elegant proof than I was expecting, especially after reading some of the earlier Erdos proofs. GPT 5.6 Sol is the real deal.
gertlabs··on Grok 4.5
Grok 4.5 is a huge step up from their next best model and now around the same performance as GLM 5.2, but it's not exactly at the frontier of the cost efficiency curve in our coding evaluations. That curve is defined by the 2 lighter GPT 5.6 models.

However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore.

Data at https://gertlabs.com/rankings?mode=oneshot_coding

gertlabs··on GPT-5.6
We have it slightly ahead of Fable in our multi-agent coding evaluations.

Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol.

Fable scores slightly lower, but with an average solution size of 12.2 KB.

Data at https://gertlabs.com/rankings?mode=oneshot_coding

gertlabs··on Claude Sonnet 5
In our coding evaluations, we found Sonnet 5 is more capable than Sonnet 4.6 (which was an underrated model itself), but is now faster and slightly cheaper.

Sonnet 5's performance is comparable to GLM 5.2 in both one-shot coding and agentic ability. However, it's about ~20% less verbose than GLM 5.2 in average code submission sizes, and uses fewer reasoning tokens, which reduces the cost gap and suggests it writes cleaner code. In practice, Sonnet 5 ends up being 40% more expensive and ~2x faster than GLM 5.2 in our evaluations (not 300% more expensive as the per-token pricing would suggest). Granted, GLM 5.2 is an extremely reasoning heavy model.

Overall, it's a solid release that gives Anthropic some standing in the price-conscious inference market.

Data at https://gertlabs.com/rankings

gertlabs··on GLM 5.2 beats Claude in our benchmarks
This is something we omit for a few reasons but it's probably the biggest blind spot in our evaluations; we opt-in to auto-reasoning/adaptive reasoning or max thinking token budgets where supported (supported by most models now), but when an explicit reasoning level is required, we fall back to High reasoning. In practice, we've found most models scale High-><whatever marketing term is max reasoning> pretty consistently, but if one vendor started throwing 10x the resources into max reasoning and they didn't support auto-reasoning, they would be unfairly penalized in our evaluations.
gertlabs··on GLM 5.2 beats Claude in our benchmarks
It would have made things easier for us if Sonnet 4.6 scored lower, but it's a great model and the data is real.

It doesn't have a higher capability score than Fable, though. We break our coding evaluations into 2 parts, and "one-shot coding" makes up part of the index, where Fable significantly outperforms every other model, which is why it's ranked at the top despite Sonnet 4.6 having a slightly higher median (and lower average) in long-horizon agentic workloads. One-shot coding tends to be the most correlated with other companies' model cards, whereas agentic coding is partly about how well a model can adapt to a custom harness. Fable also refused some tasks.

Data at https://gertlabs.com/rankings?ow=1&mode=oneshot_coding

gertlabs··on GLM 5.2 beats Claude in our benchmarks
We've spent some time trying to understand this anomaly, even re-running Sonnet 4.6 through our evaluations to see if that would bring down its scores... and it didn't. I don't know what they did differently, but it's basically Opus 4.6 with more temperature variability (some great responses, some less great, with an approximately frontier median response in agentic work specifically). It is smart, methodical and excellent at tool calling in our custom environments.

We now use Sonnet 4.6 for a number of internal use cases we wouldn't have considered otherwise.

gertlabs··on GLM 5.2 beats Claude in our benchmarks
We use a rotating pool of ~100 games for the coding parts of the benchmark, and are scored objectively based on ratings similar to Elo. Models write code submissions to interact with the environment, then are evaluated in large batches against other submissions.

We test 11 popular/interesting languages (you can see the Languages chart to filter), but not Elixir -- although other evaluations have found that many LLMs solve more problems when working with Elixir [0]. Why models write code well in some languages over others seems to go beyond pre-training data (Python scores quite low for most models) and we don't fully understand it.

[0] https://elixirforum.com/t/llm-coding-benchmark-by-language/7...

gertlabs··on GLM 5.2 beats Claude in our benchmarks
It's 100% due to tool use -- Flash adapts much better to our custom harness with tool names that are not identical to what models were likely trained on. DeepSeek V4 Pro performs much worse in that aspect than almost all other recent releases, for whatever reason.
gertlabs··on GLM 5.2 beats Claude in our benchmarks
Scroll to the bottom for the methodology (sorry, this should be linkable)
Page 1 of 5Next →