Second, compare to older versions of competitor s models.
Still does not look good? Compare to own previous models.
To be fair, seems more correct to compare against similar strength models if your main edge is pricing.
“Good” benchmarks to gauge development skills at the moment seem to be:
DeepSWE [0] by Datacurve
FrontierCode [1] by Cognition
And then there’s TerminalBench, which I’m certain has been saturated in post-training to no end, so I wouldn’t think of it as a gold standard anymore.
But yeah, in general, it’s not going to get easier knowing which benchmarks are actually measuring “frontier” capability, and which are just getting results inflated by way of time/token budget [2], ever again.
[0] https://deepswe.datacurve.ai