It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
Last week it felt like Opus 4.8 was moving the Pro "usage" meter very quickly. Today, pre-announcement, Opus 4.8 Medium felt like there was less meter-use per minute. And post-announcement, Opus 5 Medium also feels more efficient, allowing more work in the 5-hour window.
Completely subjective, of course.
"Your weekly Claude Code limit is 50% higher through August 19"
> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.
They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?