2,746 karma · joined January 29, 2013
SWE bench from ~30-40% to ~70-80% this year
https://arxiv.org/html/2410.11840v1#:~:text=Scaling%20laws%2....
"Raising visibility on this note we added to address ARC "tuned" confusion:
> OpenAI shared they trained the o3 we tested on 75% of the Public Training set.
This is the explicit purpose of the training set. It is designed to expose a system to the core knowledge priors needed to beat the much harder eval set.
The idea is each training task shows you an isolated single prior. And the eval set requires you to recombine and abstract from those priors on the fly. Broadly, the eval tasks require utilizing 3-5 priors.
The eval sets are extremely resistant to just "memorizing" the training set. This is why o3 is impressive." https://x.com/mikeknoop/status/1870583471892226343
"The Turk was not a real machine, but a mechanical illusion. There was a person inside the machine working the controls. With a skilled chess player hidden inside the box, the Turk won most of the games. It played and won games against many people including Napoleon Bonaparte and Benjamin Franklin"
https://simple.wikipedia.org/wiki/The_Turk#:~:text=The%20Tur....
Incorrect round
I think OSS is best for this category right now.
Sergey shouldn't be pair programming right now he should be doing the hard work with Sundar of bringing some semblance of coherence to their offerings but it seems like the organizational dynamics are too challenging for them at this scale.
Gwern(or anyone else) do you have any resources on this?