Christ this forum has become intellectually dishonest.
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
[0] https://status.claude.com/incidents/72f99lh1cj2c
[1] https://anthropic.com/engineering/a-postmortem-of-three-rece...
[2] https://texxr.com/handle/claudedevs
source: https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2
There are no bugs, just happy little accidents.
That just means they don’t reduce model quality for those reasons.
They didn’t mention other reason, for example, “Make more money”.