Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?
I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.
And, while you might be billed fewer tokens as a result (because the lower thinking would result in less investigatory work), you might not know this is happening, and know to dial up effort accordingly - you'd simply get a worse work product. And certainly, Anthropic's incentive for anyone on a subscription is to push this as aggressively as they can, so people use less of that subscription.
Sadly, I'd also expect that the OP's benchmark will be detected as a test of model capabilities, and thus be given a high classification so that this strategy remains undetected.
fuck