"They called me bubble boy..." - some dude at Deutsche.
"They called me bubble boy..." - some dude at Deutsche.
Probably very expensive to run of course, probably ridiculously so, but they were able to solve really difficult maths problems.
It's not real, they are cheating on benchmarks. (Just like the previous many times this was announced.)
Reasoning models didn't even exist at the time, LLMs were struggling a lot with math at the time, now it's completely different with SOTA models, there have been massive improvements since gpt4.
My point is that even if things are pleatuing, a lot of these advancements are done in step change fashion. All it takes is one or two good insights to make massive leaps, and just because things are plateauing now, it's a bad predictor for how things will be in the future.