The model did fine.
Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.
In this moment, andai was enlightened.
The model did fine.
Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.
In this moment, andai was enlightened.
And small models have the advantage that they tend to be much faster. This is an interesting factor because, well speed is always a nice-to-have, but if you go from 10 seconds per turn to one second per turn, the activity actually becomes interactive.
It becomes a fundamentally different way of working. You stay active and engaged the whole time. And also because you are "driving", your mental model does not be synchronized from the code base. So you don't need to spend extra time later catching up.
Best could mean different things to different people.
https://cognition.com/blog/frontier-code has a dropdown selector for Tokens, Cost, Time, Agent Steps, and more.
---
As a side note, these two benchmarks appear to be more sensitive at distinguishing supposedly frontier models from each other. But they themselves cannot agree on which is better!
One argues that the other has a flawed methodology (and makes a fair case). However it might also just be that the frontier is a little jagged, and sometimes one model will do better than another.
That's been my experience anyway. When I have a very important task, I always make sure to get a second opinion: I get them both to give it a shot, and then to critique each other's solutions. You can actually apply this at every stage (the review, the planning, the implementation), if you have the patience for it.
(Would be nice if there were a way to automate this process. Maybe with one of the higher level meta-harnesses that runs Claude and Codex as subprocesses. But I haven't looked into that yet...)
At any rate, combinations of models have always been found to do significantly better than a single one (e.g. see Model Alloys, Model Fusion, etc.). So make sure to use them, when appropriate!
Hope this helps.