You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?