11 karma · joined July 15, 2020
What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close.
That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.
I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.
- METR's time horizon
- TrackingAI's offline cognitive test
- Humanity's Last Exam, and
- ARC-AGI-2
that have lasted longer than 2 years (though ARC-AGI-2 is now saturated).
When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).
So, that's how you get a break even at 2.6 years...
Jailbreak approaches like "Bad Likert Judge" ( https://unit42.paloaltonetworks.com/multi-turn-technique-jai... ) and similar persuasive techniques (see https://xthemadgenius.medium.com/how-persuasion-techniques-c... ) move the text domain to more policy, analysis, or scientific papers, where deeper analysis, discussion, and compliance is the norm.
So I'm curious about the extremes (variance) of success with threatening vs. polite discussion, but I haven't seen direct research on that.