https://x.com/natolambert/status/1946569475396120653
OAI announced early, probably we will hear announcement from Google soon.
https://x.com/natolambert/status/1946569475396120653
OAI announced early, probably we will hear announcement from Google soon.
The key difference is that they claim to have not used any verifiers.
If you mean pure as in there’s not additional training beyond the pretraining, I don’t think any model has been pure since gpt-3.5.
Big if true. Setting up an RL loop for training on math problems seems significantly easier than many other reasoning domains. Much easier to verify correctness of a proof than to verify correctness (what would this even mean?) for a short story.
There's a comment on this twitter thread saying the Google model was using Lean, while IIUC the OpenAI one was pure LLM reasoning (no tools). Anyone have any corroboration?
In a sense it's kinda irrelevant, I care much more about the concrete things AI can achieve, than the how. But at the same time it's very informative to see the limits of specific techniques expand.