I guess there is no way this can happen without benchmark being part of the training data??
It seems I was wrong. American AI companies might actually be benchmaxxing harder.
All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.
Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)
Data at https://gertlabs.com/rankings
Now I am not really specifically accusing Anthropic of anything here, I'm just saying their behavior is suspicious. Since you tested Fable, they wouldn't even have to lie to have optimized for your specific benchmarks, since they absolutely had permission to read your sessions if they wanted to. But obviously, that's only the situation if we take them at their word. Personally I would be a bit surprised if they just flat out were lying and secretly retaining data they say they are not, but not that surprised. The penalties for doing this are probably worth the rewards if it keeps them super far ahead in the benchmarks for years without anyone catching on.
(In actuality though, even if they really were trying to sneakily grab samples of benchmark tests via their Fable data retention rules, I don't really suspect there would've been very much time to optimize Opus 5 on it. So consider me bothered.)
"...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."
They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.
It's more likely that the training data was contaminated with the benchmark data.
They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .
doesn’t appear to be a very strong argument
[1]: Recent example: https://www.anthropic.com/research/global-workspace
The best way to know that christmas trees are typically purchased in december isn't to access historical financial tables and run the entire analysis yourself but to just "know it" from the articles you read attesting to the fact.
If you've found data that is the answer to the question, be it a random user question or a benchmark, the softwares' goal is to produce the correct solution and it is cheaper to retrieve from storage than to compute.
It's the same with coding. The agent usually isn't really thinking about the problem from first principles, it's just giving you the answers it's already found in instances where someone else asked the same question.
It seems to me that if you're really testing for reasoning capability, just as if you were an instructor administering a test, you'll need to change the test from run to run in order to make sure the agent/student isn't just copying old tests.