It seems like there’s a week by week and sometimes day by day change in performance when on a subscription plan using their harnesses.
It seems to do really this you would need to crowdsource it -- users individually give the lab access to a body of subscriptions normally used by average people, and the lab occasionally runs some masked version of the task through on diverse accounts.
might be the case for you as well
There's clearly random variation, but it also shows each model release is just genuinely better. ( With the exception of a heavily degraded week of opus 4.7, which was acknowledged as a problem at the time. )
There's a psychology of getting used to models after being wowed by the new performance. It sets in as your new baseline expectations, and then when it doesn't deliver, it's felt more acutely. When it does deliver, it's just meeting expectations.
Then a new better model comes along and it's a step up again, another wow moment for a week or two until expectations adjust to meet the new baseline.
Even if they use a subscription account, surely Anthropic can tell which one it is.
Users are not reliable or consistent model evaluators. Users adapt to models - the moment they get a model that performs better, their expectations rise, their tasks get harder and their prompts get shorter.
"They made the model worse" is PEBKAC in 9 cases out of 10.
Then that 1 out of 10 case where the model was actually made worse (whether intentionally or by mistake) would stand out instead of being swallowed by the noise floor.