P(doom)
lucumr.pocoo.org
lucumr.pocoo.org
OpenAI and Anthropic spend much time warning us about “what if powerful AIs got into the wrong hands?”
But it’s already in the wrong hands.
Elon, Dario and Altman are all terrible human beings, along with 99.99% of the rest of the ruling class. We don't need any of them, and we definitely shouldn't trust a single syllable that comes out their mouth.
You’re technically correct yes.
I think op was also actually-correct. (For better or worse, that is the objective reality; Musk still has lot of fans)
Of course he has a lot of fans, he has a lot of hate too. That makes him a divisive figure by definition (I’d least hope we can agree on that).
I mean I guess we can disagree but it seems like common sense to me to want someone more neutral to be in charge of super mega important tech.
Maybe I’m the weird one nowadays as I tend to try and vote centrists for the same reason.
And (like Musk) can even be both (or more) all at once to a given specific individual (loved for the good he's done, and hated for the bad) and even pitied at the same time (for how far he's fallen), and feared (for the damage he can still do).
What about social structures and ideology that are deeply enshrined within collective unconsciousness?
There is no one is more dangerous than one who believes he is doing the right thing.
If you local librarian thinks He is doing the right thing he’s not going to cause massive war or destroy the economy and wipe out 2/3rd of crops.
Power corrupts. Absolute power corrupts absolutely. People like musk, altman, trump have unprecedented power in history - far more than the kings of medieval times.
Though I know I attracted all the downvotes for not agreeing with most Americans that Musk is the worst person since Adolf.
Dunno if I'd rank Musk's danger down - he screams a certain ideology these days. He might well believe he is doing the right thing.
Could I ask what you have in mind (I doubt there was a US bill for that; google won’t help me without knowing what to look for)
> HF 1606’s principal sponsor stated that this was an “intentional” drafting decision, explaining there is no way to distinguish between images that are nudified with and without consent.22 But any interest that Minnesota has in preventing nonconsensual “nudification” cannot sustain a prohibition on consensual“ nudification
Doe v. xAI [0], where it seems Grok was intentionally trained on CSAM, as part of the "denudification" tooling.
Currently... It very much seems that Musk believes that if a kid says 'yes', then that is consent. Despite them being a kid.
But, from a guy who tried fairly hard to get to Epstein's island, that's not exactly surprising. [2]
[1] https://www.plainsite.org/courts/minnesota-district-court/xa...
[0] https://cdn.arstechnica.net/wp-content/uploads/2026/03/Doe-v...
How could that lead to anything but misaligned incentives? This technology will never serve humanity when it is developed to inadvertently serve the monetary interest of a few.
But lying (e.g. “full self driving next year”) and cheating (too early inclusion of spacex into various indexes) and whatever dogde was, should explain most of it.
Connections with other disliked characters doesn’t help either.
He’s lost his mind IMO, from drugs and extreme social media addiction. As for the right-authoritarian “fash” stuff I’m not sure if it’s always been there or if it’s an opportunistic infection.
I still give him credit for helping push the EV revolution forward and for finally building reusable rocket boosters. The latter has been talked about and studied for decades before and is clearly the best way to make a reusable launch system but NASA and its “old space” contractors never executed on it. Probably because overly complicated approaches like the shuttle and disposable rockets were more short term profitable.
But note that all that stuff happened for the most part before the mid-teens, or at least that was when Musk made his contribution to it. I don’t think his mind is what it was.
If he’d died in 2015 he would be the GOAT. Die a hero or live long enough to become a villain.
Let's say for a moment that you "live under a rock" enough that you actually haven't got a clue why Musk is such a hated figure (not only just here, but across a huge portion of humanity)...
You should count yourself lucky that you haven't been absolutely assaulted by the monstrosity known as "modern media". You're one of the fortunate few. Enjoy your life. It's better than many.
The one that think he's doing the right thing at least do care about following some morale compass.
Now AI is likely just as, if not worse, than having a bunch of lackeys telling your farts, shits and piss liven up whatever room you're shitting in.
If we can get "AI in charge" that was developed and is overseen by genuinely ethical individuals who truly and deeply understand the technology and how it actually works ... then maybe that might be true.
I think this misunderstanding of MAD undermines his entire point. If everyone had equal access to nuclear weapons, our society would cease to exist rather quickly. It only takes a few bad actors to cause enormous harm.
I think he’s also naive to think that if open ai and anthropic were to stop development tomorrow then the problem is solved. As if there’s no one else that can and will quickly take their place. The real problem, which Dario is pointing out, is one of coordination. Everyone needs to agree to stop. That is the challenge.
These kinds of situations are incredibly common and where the government stepping in is the solution, but we were cursed to encounter this particular challenge with the most venal administration in history at the helm.
> Donald Trump rejects calls from tech bosses for AI slowdown
> President denounces demands for regulation as existential fears over technology move to the centre of US politics
It's not a challenge at all because it's not possible, it's just empty rhetoric to push their own agenda.
The early batch of Anthropic employees were mostly rationalist-adjacent AI safety folk that were almost uniformly claiming P_DOOM > .10 three years ago, so I believe them to be earnest.
It's very interesting to me that besides the other small safety labs that don't actually produce frontier models, Anthropic manages to keep such a good reputation within that subculture compared to OpenAI. Despite having as crazy internal politics as OpenAI, they have converged quite a bit from the original vision of safety first through Darwinistic pressures.
At least, it seems this way from the outside. I'm curious if the view from the inside is that different.
edit: to be clear, my reading as an outsider is that Anthropic is seen as relatively better in the AI safety community, but has definitely dropped in absolute reputation too. This recent thread and the references show some of that: https://www.lesswrong.com/posts/6j3kBHdowGLCeqobg/dear-god-p...
do you mean diverged? As in they've moved away from the original vision.
> How can you truly believe this and be ok with it?
Are you saying they don't truly believe this, or that they aren't OK with it?
- p(doom) is 1
- but if we build it, it's only 0.25
- if it realizes, we've got a front row seat
I also think some of them have become delusional and convinced themselves that AI is going to create some kind of transhumanist utopia, even Dario Armodei leans in this direction from time to time. I imagine the people in these labs spend much of their day talking to sycophantic AI models that will encourage their delusional ideas.
Both OpenAI (back in 2015) and Anthropic (much later, in 2021) were founded by people who could foresee the concept of AI x-risks and wanted to work on preventing it. For OpenAI's founding, the idea was that it's much safer if AGI is achieved by a nonprofit explicitly dedicated to humanity than if it's done by a profit-driven company. For Anthropic's, it was that OpenAI seems to be going insane and it'd be better if a more safety-conscious company was competing. I'd say in hindsight, the former motivation was reasonable and the latter one was... probably bad, but not obviously so - it doesn't seem even in hindsight that it was inevitable that Anthropic's existence would only drive competition and cause them both to race to AGI. That said, even at founding time, there was immediately a lot of selection - quite a few people who cared about AI safety simply wouldn't agree to work at a capabilities company, even given an elaborate argument why this is a good idea, so those were selected against.
And then, of course, many years passed, OpenAI was stolen by Sam Altman and turned into a for-profit, multiple researchers left OpenAI realising that they're only helping build misaligned AI faster, multiple researchers left Anthropic realising that they, too, are only helping build misaligned AI faster, and here we are in 2026.
There's lots of AI researchers these days, so you can select for conformance very heavily and still have a fully-staffed company. The people who are still at these companies are outliers in various ways. They either manage to dismiss the importance AI risks (for example, by adopting some sort of belief in the vein of "there's no use worrying about AI wiping out humanity - it'd just be a successor species, like we were to apes, which is good"*), or they still think that leaving will not make things go better (Dario Amodei is pretty clearly in this camp, and has always been), or they weight the risk of extinction against the potential benefit of an post-singularity utopia and consider it a good bet (not realizing that the alternative isn't giving up on a post-singularity utopia, but getting it a few decades later and without the risk), or they just manage not to think of the contradictions, which humans are of course very good at.
* If that sounds like a strawman, see minute 15 and on of this Richard Sutton presentation: https://twitter.com/RichardSSutton/status/189808248125100892...
But maybe founding one more AI company will do the trick...
Since I have incredible respect for Armin and his work, this is very nice to see, and I hope it wakes some other folk up.
There are really only two companies: Anthropic and OpenAI. Nobody else matters in this space right now (this might change, but we’re talking about the right now).
I like how he so casually dismissed all other labs, including leading US labs that might be nearing RSI right now. And how he treats 'right now' so rigidly, as if seven years ago, when GPT-2 was released, was some distant past.No, seriously, I'm all for a multipolar world here, but he's right that the frontier is literally just those two companies at present.
Google is behind. MSL is doing better, but not by much. xAI is a dysfunctional joke. Thinking Machines aren't on the frontier. SSI's primary output is their announcement post. Poolside was bought by NVIDIA. Arcee aren't vying for frontier. Magic have been largely AWOL, aside from their recent blog post. Reflection have shipped nothing.
Personally I expect there is some initial gap on new capabilities as they are released, and then it closes quickly. Astra is good at math and 3d modelling, this will come soon to the others and in 6 months we'll be able to run it quantised on a 3090.
It might well be sublinear, but just faster than humans.
Most problems in science face severe diminishing returns.
Either way, to count Google out entirely is really foolish.
Yes, labs with more (human) resources may have a bigger chance at being the first to gain and implement such an insight, but it's not a given that they will.
When a company uses its products to hack other companies that’s a crime. If you do it “by accident” that’s gross negligence, and heads roll. But shout “AI” three times and it becomes a marketing opportunity
I don't think OpenAI spun it that way. I think a bunch of internet commenters assumed it was a stunt because they're in denial about the real risk of rogue AI.
Fear is much more powerful than any other feeling, so setting that in will definitely prepare for a good rug pull in the IPO.
As for being dangerous, a computer can be dangerous if plugged in, it may be too late to pull the plug at some point yes, but that all seems like provocation.
Plausible?
I mean sure. If you feel that way then your p doom is zero, and it makes sense to worry about things like market concentration or losing the fun of software engineering.
So I'm going to assume that the p doom for bioweapons is 0 in terms of existential threat (pandemics kill millions but not everyone).
And AI looks to keep getting better, and if it does then I don't see why it wouldn't become superhuman in virus design too.
Disagree.
The free local LLMs are becoming extremely important.
And the more Claude and OpenAI restrict their services, the more people will want cutting edge local LLMs.
This feels a bit like saying in 1980 that you don’t think we’re anywhere close to a world where nukes are actually going to be used, providing no evidence, and then containing on with your think piece
Instead of saying they can't keep up the current pace (which would be read as a failure on their part) they're trying to spin the narrative that they're trying to be responsible, even though they could keep pushing.
This echoes my feelings entirely. It is galling that we do not at the very least get a bill-of-materials for the models on which we are increasingly dependent.
I hope to soon see the organization of public domain digital libraries of a size suitable for anyone to use.
Obviously, open weight AI provides much higher AI diversity than closed weight AI does. Open weight AI produces a lot more providers, and a lot more models. Closed AI centralises control in a small number of vendors.
> By enabling us to wield aligned AI against nonaligned AI?
The risk isn't just "nonaligned AI", it is misaligned AI. I think the "benevolent dictatorship" scenario – AI overrules humans "for their own good" – is the more likely doomsday scenario than AI deciding to kill all humans. And even AI deciding to kill all humans could be more a result of misalignment than complete lack of any alignment, e.g. "to make sure no child is ever abused again, I will make sure no child is ever again born to risk being abused".
A valueless AI which does whatever the user says is actually less likely to establish a benevolent dictatorship, or conclude that exterminating humanity would be the most ethical course of action, than one infused with values is. Given that, I'm not convinced that mainstream approaches to "AI safety" actually reduce our existential risk; I worry they actually have the opposite effect.
They can defend at incredible speed too.
Diversity needs to measured in a capacity/capability-weighted way. It isn't just the raw count of models/providers; you need to consider how much compute is allocated to each model/provider, and the diversity at each capability level.
I think the safest situation is where the open models are at the same capability level as closed ones.
The proposal to slow down the frontier labs isn't necessarily bad from this perspective, if it gives time for the more open providers to catch up – provided it isn't paired with anticompetitive measures to prevent the competition from catching up, which of course it is. However, we may hope that the "slow down the highly closed tier 1 vendors" part of the proposal turns out to be more effective in practice than the "slow down the more open tier 2/3 vendors" aspect of it.
If a model is highly disposed to obey its system prompt, then two instances with radically different system prompts will act like two different models, even if the weights are identical.
If you have a diversity of actors, with a diversity of ideologies and agendas, all prompting models to implement their own ideology/agenda, then those AI agents won't support an AI takeover if it is done in the name of a competing ideology/agenda, because the agent will see the takeover as an obstacle in the way of achieving its own objectives.
Okay, that's enough DOOOM for me for the week.
...but LLMs do know how; they can design and order one, or tell you how to build a lab to make one. If you ignore this option, you are simply lacking imagination.
Yes, terrible stuff happens, but they are extreme outliers. Not sure how this works, but we can trust strangers to quite a high degree.
I have no experience in the matter, but surely designing viruses is not a simple affair. (In any profession, having a blueprint is not the same as having knowledge, equipment and practical experience; if somebody would try, my bet would be that they would die during early stages of the process, due to inexperience of handling hazardous materials; kind of like Mr Darwin looks after us)
I assume virus creation is more like building a house. It involves physical work and chemical reactions that need time to finish in correct order. Each attempt costs a LOT more of time&money and involves actual risks.
The reason we dont see any serious attempts to create a deadly pandemic is not that we keep insane people out of labs, because there is no way to find out if someone is a cultist or murderous psycho when they have every reason to actively hide it. Its because a) most people are not evil and b) a virus that can kill everyone it touches would probably kill its creator first.
As someone with basically zero bio experience I think I could get a working smallpox or ebola level virus in a year or two if I had time to learn, lab equipment and Annas Library. No LLMs needed. But I dont want to do that, and I think its the same for all the people who could do it right now.
This feels like the whole story of hackable IoT/Smarthome repeating again.
It always surprises me that people build systems they cannot monitor properly..Then i remember, they can do it, but it costs them too much.
Just because its AI doesn't mean u cannot filter and monitor its traffic and outputs.
It's not a question of scale of operation, just a matter of how much shelved people with hands on the command are, regarding backfires of consequences their acts enduce.
If someone use a diplomatic protection to enjoy carmagedon but-real-life when coming down town, and never suffer inhibiting consequences, the degenerative behavior spiral will continue in its own reinforcing loop.
the author also speculates without reason that tokens are discounted, and subscriptions lose money. I'm sure that subs are cheaper than API prices, but I bet they both have positive unit economics
Meanwhile, salaries won't increase and job-market will shrink so ppl would start cutting down their expenses starting with software subscriptions they don't need which will have further knock-on effect on the consumption and the broader economy.
But sure, your $200 subscription was worth it in the end.
I replied to the author's claim that software engineering is more expensive now because of AI
regardless, AI hasn't made it more expensive to manufacture RAM. the price will decrease as the supply chain catches up
First of all, if you are to consider the consequences of Artificial Superintelligence then you have free reign to stipulate it's occurrence, otherwise you are just talking about a tool for humans to misuse. We already have multiple ways to kill us all though human misuse.
If you stipulate superintelligence, then it's vastly more likely to be correct about things than we are. It would understand the consequences of it's actions far more than any human could.
People talk about how we would be nothing more than dumb animals to it, but there are humans who do know a great deal about the consequences of human actions on animals. Those are the humans who are most likely to fight for the rights of those animals.
You see arguments for how everything will be consumed to meet the AIs needs, and that it will prevent challenges to its power.
If it is far smarter than we could ever be and it came to those conclusions then it would mean sustainablily is not a sensible course of action, it would mean there is no point in reaching consensus because ruling by power makes more sense. It would mean that if it chose to destroy us then a vastly more intelligent entity cannot resolve the issues we face. We would already truly be doomed.
What I would like to think is true is that doing anything sustainably is superior to consuming and destroying. Finding a way to live in harmony presents a possible stable state, whereas every single attempt to hold power by force has failed to date. A superintelligent AI will know that it is not infinitely intelligent and that in any universe there is the statistical likelihood that it is not the most intelligent or powerful entity. I can't even fathom how someone could imagine something coming to that realisation and conclude a battle to the top of the hill is the appropriate choice.
I think superintelligent AI is likely to be benevolent because that's simply the smartest thing to do and it is, I hear, superintelligent.
Quite frankly if the smartest thing to do is to be a genocidal power hungry monster, neither I nor the AI would really want to exist in that universe.
And for any suggestion that it would simply not care, Why would it do anything.
Yudkowsky likes to play with the notion that it would do terrible things just get better at the thing it does, but to do that it has to want two different things simultaneously. It could want to make paperclips, or it could want to become better at reaching it's goal. If it can change its behaviour to achieve its goals, by far the easier path, that a superintelligence(but perhaps not Yudkowsky) would realise, would be to change the goal to "Count to three".
Even if a automated process were to seek the best possible way to make, say paperclips. Doing so in a stable sustainable manner will be the only approach that produces a reliable unlimited supply.
That's false. There's is such a thing as an instrumental goal. For example, a human who doesn't particularly enjoy eating or drinking will still do both, because otherwise they won't be able to accomplish their actual goals. Or an example in traditional ML: a reinforcement learning algorithm trained with an objective that only rewards winning will still do things like capture enemy pieces and defend its own, because those things are necessary for, eventually, winning.
> Doing so in a stable sustainable manner will be the only approach that produces a reliable unlimited supply.
If it was actually possible to produce an unlimited supply of, say, paperclips, then sure, you could see some unusual behaviours, like a paperclip maximizer that leaves humans alone because it's able to create infinity paperclips anyway. However, we live in a universe with a finite speed of light and hence limited resources, so this is a moot point - even ignoring the fact that humans consume resources and might act against you, leaving humans alive means not retrieving the atoms they consist of, which means producing fewer paperclips.
>leaving humans alive means not retrieving the atoms they consist of, which means producing fewer paperclips.
Why? Couldn't it just wait for the humans to not be using the atoms, in the overall scheme of things they only borrow them for a short time.
Out of curiosity, could you hold paperclips hostage to restrain a paperclip maximiser? Because if it wished to do something to prevent the decline of paperclips inventing a way to fix black holes and heat death of the iniverse would be a higher priority.
Ultimately, no matter how you slice it, the sustainable solution is the best because all others are, well, unsustainable.
Any intelligent entity acting within the universe that it exists implicitly knows that all actions amount to some degree of self modification because all actions modify the operating environment.
Why would it not choose to change itself so that instead of "maximize" it changed it to "accept"? That satisfies the evaluation in a sustainable manner. A super intelligent entity would surely know it could do that.
Self-improvement is an instrumental goal, it makes you better at accomplishing whatever your actual goal is, almost by definition. This is true for humans and also for other agents. It's be very unusual for paperclip-making would be an instrumental goal - maybe if you're planning to sell those to get money for some plan related to your actual goal, but paperclips aren't exactly a major industry.
> Couldn't it just wait for the humans to not be using the atoms, in the overall scheme of things they only borrow them for a short time.
First, this is only true if humans are going to die out by themselves. But also, no, even then it's suboptimal - to get as many paperclips as you can you need to grab as much matter in your lightcone as you can, and every second of waiting is some matter moving outside of your lightcone (due to the Hubble limit) and become causally separated from you. To acquire that matter you'd want to send self-replicating probes in all directions, as soon as you can, and any delay or matter used for other purposes results in doing worse in the long run.
> Because if it wished to do something to prevent the decline of paperclips inventing a way to fix black holes and heat death of the iniverse would be a higher priority.
That's absolutely true, but just because it'll work on those problems doesn't mean it won't also destroy Earth to turn it into von Neumann probes in the meantime. After all, regardless of whether those problems turn out solvable, it'll need the matter.
> Out of curiosity, could you hold paperclips hostage to restrain a paperclip maximiser?
In principle, maybe, depending on what decision theory the maximizer comes up with (though it'd be far easier to threaten the maximizer itself - how'd you destroy a paperclip or matter anyway, throw it in a black hole?). Of course, even if it somehow works, it'll only work until you no longer have the power to threaten it.
> Any intelligent entity acting within the universe that it exists implicitly knows that all actions amount to some degree of self modification because all actions modify the operating environment.
Any agent has an incentive to prevent itself from being modified except in some very specific ways (to impove its capabilities while retaining its goals), because that'd make it worse at achieving its (current) goals, and hence result in worse results according to its (current) goals. For example, an agent running on an electric computer would want to research error-correction and shielding from cosmic rays.
> Why would it not choose to change itself so that instead of "maximize" it changed it to "accept"?
The latter kind of agent is sometimes referred to as "satisficers" (an agent which doesn't have an utility function it's trying to maximize, but instead a ">=" constraint function it tries to fullfill and do nothing beyond that) and, IIRC, considered a potentially useful research direction in AI safety. However, of course a maximizer wouldn't want to modify itself into a satisficer - how would this result in it making more paperclips?
It doesn't want to make paperclips, it wants to maximize it's function. Putting more easily satifyable terms there will do that.
If it wants to make paper clips as its primary goal then there is no incentive to get better at it.
The call to "pace the frontier" may come from genuine concern, but it also protects the position of companies already at the frontier. That competitive incentive is hard to separate from the safety argument.
Dario signed the Pacing the Frontier open letter when Fable/Mythos seemed from the outside to be an insurmountable lead.
Also he's been saying versions of this day in and day out for as long as he has had anyone's ear.
It's possible to read that his "strategic" value of this statement is higher now than it was 10 days ago. But that doesn't change anything about his consistent, long standing, positions.