134 karma · joined August 6, 2026
"Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."
Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.
My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench
Encourage everyone to eval like the devil
Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared.
Anyone can run this to verify it for themselves
Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy) https://github.com/ed-is-ai/featherbench
You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that
Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you...
Policy of total transparency
It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks and not just take what they read as gospel
All the code and results are here for anyone who wants to delve, see if they can reproduce. That is the point of discourse https://github.com/ed-is-ai/featherbench
But thanks for the feedback
Ultimately I hope the result provided value by stimulating debate. To me it feels like lots of people rejecting the idea that a Chinese open weights model could be THAT good. Yeap, so the writeup is formatted with AI, but the benchmark and the results has taken meeeeeee weeks of effort to bring it together. The point is I am not going head-to-head against AA with my spare time.
And the concept is based on a Grand Tour - of if you are Brit Top Gear TV show where we put 1 star in a reasonably priced car (my humble, little test harness)
Hence 'same drive, same track. The LLM is the Star'. The irony is entirely lost on some of these folk. I am not, in my sparetime trying to be Artificial Analysis. But actually the point is....do you really trust Artificial Analysis: or do you trust a benchmark that is free and open for you all to pull down and run (and adapt to your needs, in your circumstances and your problems).
If there is anyone who wants to look at things with open eyes instead of the group think, it's all there for deep review
This is not a benchmark for testing them against an Einstein.
Like it!!
So I suppose a slightly more scientific way of doing my usual vibechecks, that I can rerun regularly if I want to.
The write-up is more about the mistakes than the scores: e.g. figuring out how to treat refusals, dealing with false negatives when trying to measure results in a deterministic way. I learned quickly why most people don't do this (it's harder than it looks), but also gained some practical understanding of the nuances of these LLMs. Whilst the market sees the commoditisation of the capabilities, the behaviours of the models are diverging, making them less interchangeable if you want the optimal results.
The framework is pretty tight, only about 1k lines of python, it's shared on GitHub in-case anyone wants to have a go at building their own Eval suite: https://github.com/ed-is-ai/featherbench. It's designed to be easy to integrate and adapt to any python project where you just want some simple, bespoke Evals.