Are Towers of Hanoi not a simple test? Or chess? A recursive algorithm that runs on my phone can outclass enormous models that cost billions to train.
A reasoning model should be able to reason about things. I am glad models are better and more useful than before but for an author to say they can’t even evaluate o3 makes me question their credibility.
https://machinelearning.apple.com/research/illusion-of-think...
AGI means the system can reason through any problem logically, even if it’s less efficient than other methods.