If indeed, as the new benchmarks suggest, this is the new "top dog" of models, why is the launch feeling a little flat?
For comparison, the Claude 4 hacker news post received > 2k upvotes https://news.ycombinator.com/item?id=44063703
For comparison, the Claude 4 hacker news post received > 2k upvotes https://news.ycombinator.com/item?id=44063703
This is a 50 minute long video, many won't bother to watch
Goodhart's Law means 2 is approximately always true.
As it happens, we also have a lot of AI benchmarks to choose from.
Unfortunately this means every model basically has a vibe score right now, as the real independent tests are rapidly saturated into the "ooh shiny" region of the graph. Even the people working on e.g. the ARC-AGI benchmark don't think their own test is the last word.