The reasoning models (o1 pro) don't have good reasoning capability when I'm asking things from them, so I don't expect o3 to be significantly better in practice even if they look good on the benchmarks.
Still, I think ARC-AGI benchmark is awesome, and the fact that they are targeting resoning is a good direction (I just think they need to research more techniques / theories).