I.e. the AI 2027 guys were memed for being AI lunatics on a lot here and they have been pretty on the money in terms of pace of progress accelerating/ gov moving towards nationalization/ coding agents
I.e. the AI 2027 guys were memed for being AI lunatics on a lot here and they have been pretty on the money in terms of pace of progress accelerating/ gov moving towards nationalization/ coding agents
[1] - https://garymarcus.substack.com/p/breaking-the-ai-2027-dooms...
Gary Marcus’ takes have aged really poorly over time ironically he is the “AI expert” who has to constantly move the goal posts.
From my perspective there’s very slow but very real progress happening in the AI space. I see people making wild predictions in both directions, but in terms of actual unsupervised utility there’s definitely progress abet wildly slower than most hype.
"The bet of using AI to speed up AI research is starting to pay off.
OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D. Overall, they are making algorithmic progress 50% faster than they would without AI assistants—and more importantly, faster than their competitors. The AI R&D progress multiplier: what do we mean by 50% faster algorithmic progress?
Several competing publicly released AIs now match or exceed Agent-0, including an open-weights model. OpenBrain responds by releasing Agent-1, which is more capable and reliable.28
People naturally try to compare Agent-1 to humans, but it has a very different skill profile. It knows more facts than any human, knows practically every programming language, and can solve well-specified coding problems extremely quickly. On the other hand, Agent-1 is bad at even simple long-horizon tasks, like beating video games it hasn’t played before. Still, the common workday is eight hours, and a day’s work can usually be separated into smaller chunks; you could think of Agent-1 as a scatterbrained employee who thrives under careful management.29 Savvy people find ways to automate routine parts of their jobs.30
OpenBrain’s executives turn consideration to an implication of automating AI R&D: security has become more important. In early 2025, the worst-case scenario was leaked algorithmic secrets; now, if China steals Agent-1’s weights, they could increase their research speed by nearly 50%.31 OpenBrain’s security level is typical of a fast-growing ~3,000 person tech company, secure only against low-priority attacks from capable cyber groups (RAND’s SL2).32 They are working hard to protect their weights and secrets from insider threats and top cybercrime syndicates (SL3),33 but defense against nation states (SL4&5) is barely on the horizon."
That's precisely where we are.
This is eerie. It's like a time traveler. The only delta is Anthropic is in the role of OpenAI.
https://asteriskmag.substack.com/p/before-he-wrote-ai-2027-h...
https://www.lesswrong.com/posts/6Xgy6CAf2jqHhynHL/what-2026-...
Anyone who wants to dismiss the LessWrong / X-Risk / "doomers" should link their accurate predictions from 2021.
That seems to me to be the most concrete and least obvious prediction in the quoted text.
I don't think that's happening. If that were generally accepted as true I would expect OpenAI to be unable to successfully IPO.
3 * N < 42 * N
42 - 3 < N * (42 - 3)
It helps to know layer you're working on.
People seem to make mistake of thinking how good LLMs are around tasks that they are familiar with and extrapolating it to whole population.
It's good mental exercise to think about how little you can do compared to expert on tasks you never thought of working.
Ie. if you're programmer or know something about finance, don't think how much it enables you to do better coding or investing, think instead how much it doesn't enable you to work on something you don't know like maybe molecular biology or visual special effects – it's all there but it's much better multiplier for people who do know their shit.
Knowing layer you're working on helps a lot, it gets multiplied.
Knowing programming is becoming more fundamental skill than ever before as it lies at the foundation of almost everything else.
I’ve heard people say older models can’t do X, when I used that way etc. I suspect people are applying their own learning curve as part of their assessment of progress, you get better at writing prompts and it feels like the model improved.
Which is why I’m saying we need some objective metrics to judge predictions of actual capacity.
These companies aren’t just making stuff up, they really do want to improve the models, and the models really are improving.
I’m aware of multiple cases of benchmark cheating/“optimization”. So, taking benchmarks a face value seems laughable.