You've posted a link that doesn't support your statement.
https://marginlab.ai/trackers/claude-code-historical-perform...
> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).
Changing the harness can have a big impact on performance even when leaving the model completely unchanged.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.