HNHacker News
TopNewBestAskShowJobs

ed-is-ai

134 karma · joined August 6, 2026

submissionscomments
ed-is-ai··on GPT-6 Astra
Pretty nice pelicans!
ed-is-ai··on Dyson CameraJet electric toothbrush
ok first thoughts - this is a cool, but ludicrously expensive toothbrush! Anyone who buys one is just flaunt their wealth when $2 gets you a manual version
ed-is-ai··on Claude Fable 5.1 and Claude Mythos 5.1
The big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, not ours

"Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.

Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.

My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench

Encourage everyone to eval like the devil

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Noted. Feedback received
ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Indeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases OpenAI, Anthropic models should not be the default choice any longer.

Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared.

Anyone can run this to verify it for themselves

https://github.com/ed-is-ai/featherbench

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
I had to google astroturfing. I liked the due diligence... I am a hacker news infant, so all my history is about this.

Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy) https://github.com/ed-is-ai/featherbench

You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Impolite
ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
It's good to check the facts.

Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you...

Policy of total transparency

https://github.com/ed-is-ai/featherbench

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Thanks for point out - I updated the benchmark today to include. It is every bit as good as everyone says it is. The intelligence / $ is something else...
ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
That's pretty cool. All available already in OpenCode if you're a developer, see for yourself!
ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Thanks for the feedback.

It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks and not just take what they read as gospel

All the code and results are here for anyone who wants to delve, see if they can reproduce. That is the point of discourse https://github.com/ed-is-ai/featherbench

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
https://github.com/ed-is-ai/featherbench
ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
pcwelder -https://github.com/ed-is-ai/featherbench all the data is here if you care to look

But thanks for the feedback

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Hmmm, as the author: the feedback is interesting indeed and something to take onboard...

Ultimately I hope the result provided value by stimulating debate. To me it feels like lots of people rejecting the idea that a Chinese open weights model could be THAT good. Yeap, so the writeup is formatted with AI, but the benchmark and the results has taken meeeeeee weeks of effort to bring it together. The point is I am not going head-to-head against AA with my spare time.

And the concept is based on a Grand Tour - of if you are Brit Top Gear TV show where we put 1 star in a reasonably priced car (my humble, little test harness)

Hence 'same drive, same track. The LLM is the Star'. The irony is entirely lost on some of these folk. I am not, in my sparetime trying to be Artificial Analysis. But actually the point is....do you really trust Artificial Analysis: or do you trust a benchmark that is free and open for you all to pull down and run (and adapt to your needs, in your circumstances and your problems).

If there is anyone who wants to look at things with open eyes instead of the group think, it's all there for deep review

ed-is-ai··on GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
Read the benchmark - it's on the basis of being 'good enough' for everyday tasks. Which is what fits most applications right.

This is not a benchmark for testing them against an Einstein.

ed-is-ai··on Building the Ed-O-Meter: Notes on Writing My Own LLM Benchmark
It looks like you're mostly 1-shotting apps and reviewing the results based on # interventions etc

Like it!!

ed-is-ai··on Building the Ed-O-Meter: Notes on Writing My Own LLM Benchmark
neat! I will check it out. I'm not really using local LLMs as dont have a big enough machine for these to run with a decent tps
ed-is-ai··on Building the Ed-O-Meter: Notes on Writing My Own LLM Benchmark
Author here. This began as a weekend project because I didn't trust the public leaderboards — as what I was experiencing seemed to differ from the professional benchmarkers. I really wanted to see how the latest LLMs worked on 'my realworld' tasks. A mix of random things like planning a holiday, getting a vegetarian recipe, debugging code, jailbreak attempts. The results are here if you want to poke at them: https://reinvently.co.uk/tools/ed-o-meter/

So I suppose a slightly more scientific way of doing my usual vibechecks, that I can rerun regularly if I want to.

The write-up is more about the mistakes than the scores: e.g. figuring out how to treat refusals, dealing with false negatives when trying to measure results in a deterministic way. I learned quickly why most people don't do this (it's harder than it looks), but also gained some practical understanding of the nuances of these LLMs. Whilst the market sees the commoditisation of the capabilities, the behaviours of the models are diverging, making them less interchangeable if you want the optimal results.

The framework is pretty tight, only about 1k lines of python, it's shared on GitHub in-case anyone wants to have a go at building their own Eval suite: https://github.com/ed-is-ai/featherbench. It's designed to be easy to integrate and adapt to any python project where you just want some simple, bespoke Evals.