Here are sites with ongoing measurements checking if a model has been nerfed - check opus 5.5-
https://www.bridgebench.ai/nerf-bench
https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/codex/
https://www.bridgebench.ai/nerf-bench
https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/codex/
Methodology section in some of these benchmarks doesn’t say if they use subscription or API. API usage may not be nerfed as much as subscription.
They must be using the same lever to “pace the frontier”. All of the best effort models from different companies have similar scores. There is no standard definition of “max” effort level.