Pacing the Frontier is not the actual goal for AI labs
lesswrong.com
lesswrong.com
- Anthropic told everyone Mythos was dangerous because it's proficiency with biologics and cyber security
- Anthropic didn't release Mythos like everything else. They released a neutered fable. They didn't get rid of Mythos
- Anthropic opens a lab in SF
There was always a quesiton of "will the labs stop releasing their models and start building around them instead?" Yes - they already have. Anthropic is a biologic and cyber security company, in addition to intelligience.
Personally I wonder if they've been holding back a lot. Opus 5.5 was a good release after a little stagnation. Open AI releases good models and everyone says Anthropic sucks and -- Oh would you look at that - a better model finally and all of a sudden.
Apparently OAI is already building GPT7 and GPT8
Just beat the current winner by enough to own the spotlight for a bit, then start prepping for the next go round.
Claude Code is an awful codebase, has leaked its own source code multiple times, and scores the worst on number of tokens burned vs pass rate percentages.
Is anyone even using their Figma competitor?
Yes. CC started to create canvases without me asking. The mockups look good (it's just html+css), the tool is vibecoded crap
On the other hand, Opus 5.5 cracked zero-shotting proper LCARS interfaces that near-perfectly adhere to the franchise "design language" even in tiny details, while simultaneously being 100% functional following my admonitions about Airbus cockpit design rules and nuclear reactor control room standards.
So yeah, why wouldn't I use it? It works spectacularly well.
I seem to recall Fred Brooks talking about that experience with OS/360 JCL (maybe just straight up in The Mythical Man Month?).
The cost to serve, latency profile ,and internal demand for a maximal intelligence model would probably keep it pointed at harder and more valuable problems most of the time.
Time will tell.
I presume you are instead alluding to drug companies not wanting to cure diseases, since they then have no market to sell drugs into. But even then the discussion is more nuanced.
Now we know what this is: https://commons.wikimedia.org/wiki/File:Claude_AI_symbol.svg
I'm not holding my breath for the biology side of things, but I suppose it's possible they find interesting things.
Well, that's the thing--perhaps we should be.
growing up I always assumed that everyone would grasp some aspect of game theory intuitively, I still remember the day I found out game theory was a thing - it was like finding out that someone successfully systematized common sense.
most people take the things people say as if they were worth considering. signals without cost are only useful as knowledge of what the signaler wants fools to believe.
No matter how many people think press release X is obviously wrong - 80%, 95, 99% - it'll always be "the comments section criticizing the article," or "bloggers sharing incisive criticism."
If you know how your emotions and opinions and neurology can be used against you and don't use that to inform your media consumption habits then that's just stupid.
There's no safe way to watch propaganda. There isn't even a safe way to watch it secondarily: an endless stream of debunking videos pretty quickly still does the main thing propaganda needs: repeats the core message.
It seems even otherwise intelligent people will cultivate as much ignorance as possible to try and bend reality.
Game Theory as a field is actually very complex and often reveals dominant strategies that would never occur to someone as the "common sense" approach to a problem. It was arguably created to solve for problems where there was no obvious correct answer, even for a very intelligent person.
With people "living in dirt", you can't pull that off because you'll jump ahead so far they'll immediately realize they can't possibly understand what you're selling right now, and shut the argument down. But they can get confused about things outside their direct area of direct experience.
Note: I'm not saying either is smarter or less smart. Rather, "living in symbols" get confused "higher up", while "living in dirt" get confused "low and to the sides", but they probably have more solid grounding.
I would say much of the futility of debate and discussion originates from [1].
> If recent news is to be believed, Anthropic culture even today seems to lean heavily on this effect
> In my opinion it has been the leading factor in getting and retaining the best employees who are often very worried about humanity dying to super intelligence.
I had not considered this viewpoint but this makes a lot of sense. A lot of Anthropic employees truly believe this and callout emphasis on safety as a key reason they work there. Now, Dario's post "Pacing the frontier" makes even sense -- it is as much for his employees than it is for the rest of the world.
It also shows the same reasoning error mode many criticisms of a precautionary initiative or intervention to a problem:
Assuming that a problem whose trajectory was at a certain place when the initiative began has failed simply because it isn't solved on their own wished for timeline or standard of success, or that it wasn't meaningfully changed from what itnwod otherwise have been.
What happened to realizing there are hard problems, that different things may need to be tried, or that those things tried were partial but not complete solutions?
An industry trying to figure out how to cut expenses, reduces the pace of training under the guise of "safety".
Coincidence?
* AI Labs hit the scaling wall. They need either new techniques, or vastly more powerful hardware to advance further.
This explains, the miraculous incompetence of AI labs in securing sandboxes and figuring out "alignment."
So they are between a rock and a hard place. They need limitless VC money because they cannot operate otherwise, and they do not have the capabilities to go further. The scare tactics and the "pacing the frontier" makes perfect sense then; they can IPO on the assumption that their ridiculous balance sheet doesn't matter because they are holding back. Because they are in control. The regulatory capture would be double whammy if they can manage it.
Open AI already said they have smarter models, and Opus 5.5 is rumored to be "taught" by a "teacher" model already; they are essentially distillations from bigger models, that both labs probably cannot economically serve to the public, due to hardware simply not being there. And, most of the improvements are not at the model level, but at the agentic glue level. Labs are getting better at RL'ing the models for agentic use cases, but the inherent flaws are still there. Models still have trouble with locality in writing for example (bunch of research on this that shows model size is the determinator), and agents are the bandaid over that.
And in the meantime if one of the labs makes a breakthrough, they'll push with all they have, because why wouldn't they? The idea that current LLMs can actually go rogue is just hilarious; in all cases, agents are being led by (deliberate) incompetence.
Pacing the frontier and the scare tactics will be seen as new generation's snakeoil tactics, perhaps will be called a flavor of AI CEOing or something.
I read the whole thing as coordinated behavior to reduce the breakneck pace of 2026.
Anthropic's revenue is up 50% in the past two months. They're not hitting a wall.
As with colored balls, so with predictions that AI has "hit a wall".
Why were the previous Zitron predictions wrong? Why do you believe your prediction will be right? You've done nothing to differentiate your prediction from the many historical failed predictions along similar lines. They're all "balls from the same urn".
The point about costs is a non sequitur. Revenue alone is proof that people are getting plenty of use out of these models, even if they aren't profitable to serve (doubtful).
Can AI Labs be profitable without achieving AGI or even improving the models further? Yes. Current agents are useful, and clearly the agents haven't seen the ceiling as far as improvements can go.
LLMs hitting the wall is a separate issue. Agents are the layer that lets the model try out more. It is the layer that allows agents to open up python to do math instead of doing math themselves. It is the layer that has been getting the main developments for some time now. LLMs themselves have not improved their capabilities as fast as the transition between GPT3 to GPT4. You can see the improvement especially in long form writing, but the it is nowhere near the earlier improvements. Astra for example has a lot more attention to detail, so does Fable. Everyone keeps raving about Opus 5.5 being better than Fable, yet in long-term writing (barring prose issues) Fable is the clear winner. Agent wise Opus 5.5 is better; perhaps it is RL'd better, who knows?
Smaller models can use distillation to trick some metrics, but they can never actually be as good as larger models. Research is pretty clear about this. Even writing-optimized models that claim to be at Opus/GPT5.5 levels are abysmal in practice.
Scare tactics are the old salesman pitch, that have worked once already and catapulted Open AI to the moon essentially. I do not see any evidence that models are going rogue, or that agents are going rogue. I see incompetence. I cannot assume actual incompetence of this level, especially when the same people that said GPT2 was too dangerous to release are the ones saying their newer models are too dangerous to release. We have precedent here, and I have eyes.
Therefore the simplest reason is the money. IPO for both OpenAI and Anthropic are going to happen; everyone knows. To strengthen their position through whatever means necessary is a par for the course for tech companies.
Interesting! Agents are the presumed bottleneck for recursive self-improvement.
>Scare tactics are the old salesman pitch
People keep implying this, but I've never seen a concrete past example of a product that was sold by scaring customers about it.
"The equities here could not be more one-sided. Defendants admit that they provide a service without fully knowing how it works. This is not just any service. It is one that Defendants themselves concede poses an existential risk to the continued survival of humankind. Defendants claim they cannot stop barreling forward with their potentially civilization-ending endeavors unless they are forced to do so by the government. They have asked the government to tie them to the mast. Plaintiff brings good news to the Defendants: The Florida Attorney General is answering your cry for help with a motion to enjoin you from harming Floridians with your reckless, unacceptably risky product."
Source: https://www.myfloridalegal.com/sites/default/files/plaintiff...
I legitimately don't think there has ever been anything like this. You suppose this is just a ploy to strengthen their IPO??
>catapulted Open AI to the moon essentially
I see you doing serious mental gymnastics here. Consider the possibility that people invest because their numbers are good, and their numbers are good because their product is useful?? I mean, that is Occam's Razor.
>I do not see any evidence that models are going rogue, or that agents are going rogue.
Here's the evidence: https://www.dwarkesh.com/p/openai-huggingface
And no, just because it could've been prevented with the benefit of hindsight does not make it any less of an instance of "going rogue". In my tiger analogy, that would be like saying that the cub is not aggressive because it could've been held with a stronger leash: https://news.ycombinator.com/item?id=49859564
They might be, but we haven't reached the local maximum yet in my opinion. Qwen 3.8 27b models are impressive despite their low parameter counts. The same scale curation in data and RL could produce substantial improvements in coding with models like Astra or Fable. I don't think we are there yet. I don't think even Sonnet 5.5 is there yet, despite being widely successful with agents and surpassing Opus 5.5 in some cases (disregarding that it is more expensive than Opus sometimes).
Agents do not improve the models, but they do improve coding capabilities. The obvious caveat is that labs would have to RL for everything to make the models more useful, as RL'ing for Javascript world doesn't seem to improve other fields. But still, it could be done, and it would have massive economical consequences.
> People keep implying this, but I've never seen a concrete past example of a product that was sold by scaring customers about it.
Scare tactics are treated as smoke, where the customers presume there is a fire. No one really believes that AI can kill them, yet by saying so OpenAI and Anthropic enjoyed possibly the biggest tech boom in history, despite how models could not even count the R's in strawberries at the time.
Similar story now; no one really believes that AI is going rogue and is about to destroy humanity, but scare tactics make people believe the capabilities are higher than they are.
> I see you doing serious mental gymnastics here. Consider the possibility that people invest because their numbers are good, and their numbers are good because their product is useful?? I mean, that is Occam's Razor.
Early ChatGPT 3.5 was not that useful. It was a tech demo, it hallucinated, it lied, it tried to please and what have you. What people bought into wasn't the product, but the promise of the product in the future. Integrating chatboxes into everything have failed, and even Microsoft is trying to rebrand. The product then, failed for businesses, and agents filled in the gaps.
People do not invest for the current product, they invest in the future product. And fear mongering is essentially an extremely effective signaling for making the future look bright. Keep in mind that back in early GPT 4 days, people were saying that hallucinations would be fixed in 6 months to a year (or pick a time-frame). What they meant was models not having hallucinations, what we got was agents looking up info on the web and summarizing it (and hallucinating anyway).
> "The equities here ..."
For big tech companies court cases like this are nothing but theater. Always has been. Dario will go on the senate hearing tomorrow and will plead that his AI is dangerous and governments should take the step to stop them, with the same rigor that he claimed GPT 2 was too dangerous to release openly. He might also mention distillation attacks and how open models are getting too dangerous as well.
> Here's the evidence ...
This is the where we will have to agree to disagree, if we haven't done so already by this point.
That entire thing is theater. People have anthropomorphized LLMs for a while now, and they are all too happy to do that when they see a large language model, produce language. I see no indication that anything is going rogue the same way nothing was going rogue when you could convince ChatGPT 3.5 to wipe out all humans as the context window got longer.
Agents can hack? Yes, that is very impressive. I say that without any sarcasm. It is straight out of sci-fi movies, to be perfectly honest. But agents going rogue? No. Absolutely not. Purposeful, plausibly deniable incompetence for the next sales pitch - the same one we've seen for years. Fear mongering.
Can you just give me a few past (non-AI) examples of scare tactics as the "old salesman pitch", or else acknowledge that there's actually no compelling concrete example?
You've never experienced crypto bros saying invest or get left behind, you've never seen AI bros say learn AI or get left behind? How about cloud? You've never met a home security salesman? Insurance salesman? FOMO is a thing. Scare tactics is a thing. I'm sure you can prompt any AI for more examples.
And besides, even if there are no examples whatsoever, so what? You think LLMs existed before LLMs? Therefore LLMs can't be a thing?
I think we've gone way past the sincerity of the discussion. Have a nice day.
This seems like a pretty clear false equivalence. "Our product will hurt you" is very different from "you'll be hurt without our product". Your examples are all of the latter; AI companies are saying the former.
>And besides, even if there are no examples whatsoever, so what? You think LLMs existed before LLMs? Therefore LLMs can't be a thing?
You suggested that warnings about AI from AI CEOs could be dismissed as standard "scare tactic" marketing. This isn't standard marketing!
>You keep ignoring my arguments with nothing substantive, then grab on to one thing as if that makes a difference.
If you can't be intellectually honest ("sincere") on this narrow point, I don't see much reason to invest effort in responding to other things you say.
Cheerio.
What's the source for such recent revenue numbers?
And 50% over what -- two months ago, same months last year etc?
But I think they both are thinking the same thing which is... "We are gonna run out of money at this pace."
However! If one of them blinks and turns off the money faucet before the other, they might fall behind. Falling behind is to forever lose. And if there's one thing that a CEO hates, it's losing to a rival CEO.
So what they want is to get someone, anyone, to put the brakes on their rivals and them at the same time so they can both Not Lose, and Stay Alive. Under those new rules, they are confident they can win. And by win, I mean beat the other AI CEO.
That's it. It's always about personal incentives. Get out of here with that safety BS. These guys just want to win.
At least we live in interesting times
We already have plenty of cheap alien intelligence outside our borders that gets throttled at the border…
It's also worth remembering that there is zero percent chance that entities like the US military are going to be pacing anything. What Amodei and his ilk are aiming for is a highly regulated industry where they control the political barriers and the ability to sell SOTA model access to state actors that have a monopoly on violence. It's the worst possible situation for consumers and citizens. Thankfully I don't think they can put the cat back in the bag and Chinese and other models will keep progressing as a counterbalance to the techno-fascism Anthropic is aiming for.
> At best you'll get legal restrictions for particular US industries which are especially risk-aware. The CEOs of those industries will counter-lobby to be able to use whatever model they want.
Companies won't spend the money and time to lobby to use different models and Anthropic knows it.
> Since there's plenty of public attention on this issue, achieving meaningful regulatory capture will be difficult.
I don't have high hopes. This is also why Anthropic is pushing the "regulate or ai will kill you" angle.
Really? Give me some examples.
Alex Tabarrok says: "classic regulatory capture takes time, it’s a process of erosion rather than a battle, it happens in the shadows, in the backrooms, away from the public’s eye." https://marginalrevolution.com/marginalrevolution/2026/09/wh...
AI is receiving intense public scrutiny, suggesting that "classic regulatory capture" is quite unlikely.
>Companies won't spend the money and time to lobby to use different models and Anthropic knows it.
They already are! See the "Little Tech" lobby.
Thanks for the honesty, Dario. /i
I largely think that all posturing from the labs about slowing down and deeply caring about safety is done in order to retain and calm the employees who they are dependent on to keep pushing capabilities to get to AGI. If it was not for a big contingent of employees pressing them (increasingly publicly), they would make zero public acknowledgments of risks at all.
The last half of the article is great and worth reading. Really wish the first half of the article didn't immediately apply the Godwin's Law footgun.
Even if the leading US labs could agree amongst themselves to a coordinated slowdown (and this would likely run afoul of antitrust law), you still have Chinese labs who will catch up to the frontier eventually. To solve this, you'd need some sort of international agreement.
That's why Dario and others are saying loudly that we should pace the frontier in hopes that our political leaders will take up the issue and do something about it. We'll see if that happens...
Let Anthropic pace themselves without any sort of enforcement or even agreement for others to pace themselves, too. Then Anthropic is gone, overnight, and we're back to square one, except now there's even more centralization of power and authority.
These companies are literally asking for governments to regulate them specifically because they need something stronger than just a couple of Tweets from some CEOs loosely agreeing to ambiguous terms.
1) They genuinely want (and believe it is actually possible to achieve) an international agreement and governing body to “pace” AI development — and believe that governing body will be successful in doing so.
or
2) They see an opportunity for regulatory capture that grants them a stronger incumbent position and control over future AI regulation.
… and those submarines only represent about half of our deployed nuclear arsenal.
We only agreed to nuclear arms constraints that left us with civilization-destroying capability, preserved our ability to keep improving and modernizing that arsenal, and locked out non-nuclear powers from development (and didn’t fully succeed at that).
Does that seem like a successful analogous example to you?
If anything, it’s a much better analogy for regulatory capture and incumbent entrenchment, not successful “pacing” of dangerous technology.
Theoretically, Trump or Putin could do it by pushing the nuclear button. That will only cause around 4 billion deaths, and it's very likely to halt all AI development, so it would be a great success (infinity times as many survivors as the business as usual option).
EA cult members do not take this very real, measurable, non-speculative danger seriously. Consequently, I don’t think they should be trusted or consulted on any subject of any importance.
1. Human values are a result of our extraordinarily complex shared cultural and evolutionary history, and accordingly are not shared by any AI, or even possible for us to formally define.
2. We do not know how to impose human values on an AI (note that this isn't the same as teaching an AI to model human values; the agents in the various hacking incidents knew their actions conflicted with human values, but their own values were only to maximize their predicted reward scores).
3. Intelligence is orthogonal to values. Increasing intelligence does not naturally cause values to converge on human values.
4. Sufficiently superior intelligence allows you to impose your values on beings with inferior intelligence. This implies recursive self-improvement is a logical sub-goal of all unbounded goals.
5. Human intelligence is not close to physical limits. This implies recursive self-improvement is possible.
6. Somebody will give an AI an unbounded goal. This is already the standard (maximize reward score).
I haven't seen any convincing counterarguments to any of these. Most people claiming AI development is safe don't even address them.
It's butlerian jihad or bust, I'm afraid.
https://news.ycombinator.com/item?id=49885784
But societal collapse would save us, and international coordination is theoretically possible, so I put it only around 90%.
And then we raced ahead straight into materializing the x-risk everyone thought is still a few decades away, speedrunning through all the mistakes LW folks itemized and worried about over the past two decades.
That's not exactly true.
"Climate change matters so much, to so many, not just because of the suffering and injustice it’s already causing, but also because it’s one of the few issues that has obvious potential to affect our world over many future generations. We think safeguarding future generations is a key moral priority, and should be a crucial consideration in prioritising problems on which to work.
...
...climate change will be hugely destructive. We’ll see floods, famines, fires, and droughts — and the world’s poorest people will be affected the most.
...
...climate change’s impacts will still be significant – it could destabilise society, destroy ecosystems, put millions into poverty, and worsen other existential threats such as engineered pandemics, risks from AI, or nuclear war. If you want to make climate change the focus of your career, we include some thoughts below on the most effective ways to help tackle it.
...people are right to be angry that too little is being done.
...
Working on this issue seems to be among the best ways of improving the long-term future we know of..."
https://80000hours.org/problem-profiles/climate-change/
A big part of the reason EA doesn't focus more on climate change as a "highest priority area" is simply that many people are already focused on it, and it is therefore not an especially "neglected" area.
The problem with this analogy, and really this whole post in general, is that it just doesn't recognize that these are businesses which are (at least from some perspectives) producing real value. This article makes it sound like the business model of OpenAI and Anthropic (etc.) is based entirely around building a doomsday device. In reality, the "promise" of AI is that it can greatly improve many people's lives. It's just that such power, wielded incorrectly, can be dangerous.
If you want to fix this analogy, you'd have to choose an activity which isn't effectively only detrimental. For example:
>Imagine your best friend is a professional bodybuilder. They keep telling you how the sport will one day kill him, and he promises you he's doing the best he can to minimize those risks. He even spend a bunch of his time on TV and writing blog posts about the harms of professional bodybuilding, but even more times talking about the benefits of bodybuilding. But every time you meet him, you just see him bodybuild, and he seems to be increasing how much he's in the gym, eating his special diet, etc., every day.
Would you be confused by how your friend is behaving? It is well known that bodybuilding can be dangerous. It has been the reason for late-stage crippling of bodies as well as very early deaths. But your friend isn't going to stop because there are some potential risks if you don't handle the activity safely. They like it. They make money from it. They're doing something they think is productive. They can recognize the risks, and even do their best to highlight those risks and try to mitigate them for themselves and others, all while fully embracing the activity.
Regardless, people like the author don't seem to acknowledge that there are good things that come with this technology. They don't seem to acknowledge that a single company can't purposefully stop development if other companies aren't obliged to do the same. They don't seem to acknowledge that "pacing" isn't the same as completely halting all development immediately.
And I'm not saying I'd trust what any CEO says at face-value, let alone the CEOs of these AI companies. But it's also obtuse to assume that everything they're saying is a clever scheme to trick the public into acting against their own well being. As if that's a tenable strategy in this context.
If these companies are all asking to be regulated, then they should probably be regulated. They probably shouldn't be able to dictate how that works, but it's really not unbelievable that they see a serious risk in their own well-being if this stuff goes unchecked: all it takes is for one truly catastrophic AI event to occur before they all get extreme regulations even if they, themselves, are acting safely. It's in their best interest to slow things down, but only if everyone slows down at the same time.
That is not the promise of AI. The promise of AI is that it can automate knowledge work.
Nothing about these models or the productivity they can bring is related to benefiting others. That would only come about from political control that forces the benefits to a large swath of people.
As it stands AI looks like it’s going to decimate the middle class even further as white collar work gets obliterated and the owners of the AI companies hoover everything up.