I checked the developer's X account, they have written numerous posts about formal verification, so this specific claim ("without realising that said field exists") seems to be false.
906 karma · joined April 15, 2022
I checked the developer's X account, they have written numerous posts about formal verification, so this specific claim ("without realising that said field exists") seems to be false.
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
I'm not sure how many AI researchers would find this accurate. It seems to me that under conditions of ambiguity people often default to describing their preferred version of reality.
There's also a denominator problem. The mileage figure appears to be cumulative miles "as of November," while the crashes are drawn from a specific July-November window in Austin. It's not clear that those miles line up with the same geography and time period.
The sample size is tiny (nine crashes), uncertainty is huge, and the analysis doesn't distinguish between at-fault and not-at-fault incidents, or between preventable and non-preventable ones.
Also, the comparison to Waymo is stated without harmonizing crash definitions and reporting practices.
What about the 1953 CIA/MI6 coup that overthrew Iran's elected prime minister?
Why state this as absolute fact? Seems a bit lacking in epistemic humility.
https://en.wikipedia.org/wiki/List_of_dates_predicted_for_ap...
- It's not a fair match, these models have more compute and memory than humans
- Contestants weren't really elite, they're just college level programmers, not the world's best
- This doesn't matter for the real world, competitive programming is very different from regular software engineering
- It's marketing, they're just cranking up the compute to unrealistic levels to gain PR points
- It's brute force, not intelligence
"On MedXpertQA MM, GPT-5 improves reasoning and understanding scores by +29.62% and +36.18% over GPT-4o, respectively, and surpasses pre-licensed human experts by +24.23% in reasoning and +29.40% in understanding."
GPT-5 demonstrates exponential growth in task completion times:
https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com...
In 2021 Paul Christiano wrote he would update from 30% to "50% chance of hard takeoff" if we saw an IMO gold by 2025.
He thought there was an 8% chance of this happening.
Eliezer Yudkowsky said "at least 16%".
Source:
https://www.lesswrong.com/posts/sWLLdG6DWJEy3CH7n/imo-challe...
"This nearly doubles the previous commercial SOTA and tops the current Kaggle competition SOTA."
>I have repeatedly said that "can LLM reason?" was the wrong question to ask. Instead the right question is, "can they adapt to novelty?".
https://old.reddit.com/r/singularity/comments/1jl5qfs/its_ju...
It did Dragon Ball Z here:
https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...
Rick and Morty:
https://old.reddit.com/r/ChatGPT/comments/1jjtcn9/the_new_im...
South Park:
https://old.reddit.com/r/ChatGPT/comments/1jjyn5q/openais_ne...