Quality vs cost - medium is the sweet (perhaps better too!) spot.
my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward
That is just a single benchmark tho
If I need something smarter I use Fable. Medium works well and is quick. Opus 5 medium feels much better to me than Opus 4.8 medium.
yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap