Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
openai.com
openai.com
Sometimes I am baffled that sentences like this are used as headlines. I know I’ll sound dismissive… but: duh?!
> On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model.