The recent advance of reasoning in o1 preview seems to have not been widely understood by the media and even LeCun in this case. o1 preview represents a new training paradigm using reinforcement learning on the reasoning steps applied to the solution. This step allows for reasoning to be developed just like AlphaZero is able to ‘reason’ and come up with unique solutions. The reinforcement learning in o1preview means that the ‘repeating facts learned’ arguments no longer apply. Instead the AI is free to come up with its own reasoning steps that lead to correct answers and these reasoning steps are refined over time. It can continue to train and get better by repeatedly answering the same questions, the same way alphazero can play the same game multiple times.