GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
reinvently.co.uk
reinvently.co.uk
Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent
Look at the later comments, they have substance oriented discussions.
HN Mods - can you please consider a policy against such comments, it's overflowing the site and is diluting discourse and value. If people don't like an article, they can simply ignore it. These articles are reaching the top because enough people consider it of value.
But when an article comes to the top, and the topmost comment and discussion thread is an unqualified witch hunt, it's getting sick.
Per the writing, reading AI writing is like having something taste "chemical". Not very specific, but still a very recognizable and bad taste that makes it hard to enjoy and marks the thing as low quality.
And this really has to be policed. Once some tipping point is reached and too much of the HN homepage is AI slop, the site is dead.
Ultimately I hope the result provided value by stimulating debate. To me it feels like lots of people rejecting the idea that a Chinese open weights model could be THAT good. Yeap, so the writeup is formatted with AI, but the benchmark and the results has taken meeeeeee weeks of effort to bring it together. The point is I am not going head-to-head against AA with my spare time.
And the concept is based on a Grand Tour - of if you are Brit Top Gear TV show where we put 1 star in a reasonably priced car (my humble, little test harness)
Hence 'same drive, same track. The LLM is the Star'. The irony is entirely lost on some of these folk. I am not, in my sparetime trying to be Artificial Analysis. But actually the point is....do you really trust Artificial Analysis: or do you trust a benchmark that is free and open for you all to pull down and run (and adapt to your needs, in your circumstances and your problems).
If there is anyone who wants to look at things with open eyes instead of the group think, it's all there for deep review
This is not a benchmark for testing them against an Einstein.
I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.
"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...
This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
Data at https://gertlabs.com/rankings
That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.
Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.
Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.
Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.
My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench
Encourage everyone to eval like the devil
These are open weight models (GLM-5.3 soon too). You can run them on the Together AIs or Firework AIs of this world. Use OpenRouter or HF Inference Providers in between and you can effortlessly switch between models and providers.
I have been using GLM and Kimi models the last few months mixed with the latest Anthropic models and for my daily work there is barely a difference anymore (except for pricing).
Please, bloggers, write with your own voice. Don’t let an LLM do it for you.
Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy) https://github.com/ed-is-ai/featherbench
You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that
- Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers.
- X, not Y
- A, never B
- Tasteless em dashes
- Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.
But thanks for the feedback
Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you...
Policy of total transparency
It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks and not just take what they read as gospel
All the code and results are here for anyone who wants to delve, see if they can reproduce. That is the point of discourse https://github.com/ed-is-ai/featherbench
Slop