There is no credibility to any of it.
To me, this smells of XKCD 1217 "In petri dish, gun kills cancer". I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).
For all we know, the AI may indeed already be *attempting* it. They might be ineffective (hallucinated misdeeds aren't effective), or it might be why so many went from "Pause AI" to "Let's invest half a trillion on data centers".
But it doesn't actually matter what has already happened, the point is, once the AI are *competently blackmailing multibillionaires*, it is too late to do anything about it.
> I.e. idealized conditions cause specific behavior. Which isn't new for LLMs. Say a magic phrase and it will start quoting some book (usually 1984).
In normal software, such things are normally called "bugs" or "security vulnerabilities".
With LLMs, we're currently lucky that their effective morality (i.e. what they do and in response to what) seems to be roughly aligned with that of our civilization. However, they are neural networks which learned this approximation by reading the internet, so they are likely to have edge cases at least as weird and incoherent as those of random humans on the internet, and for an example of that just look at any time some person or group has demonstrated hypocrisy or double standards.
The authors acknowledge the difficulty of assessing whether the model believes it’s under evaluation or in a real deployment—and yes, belief is an anthropomorphising shorthand here. What else to call it, though? They’re making a good faith assessment of concordance between the model’s stated rationale for its actions, and the actions that it actually takes. Yes, in a simulation.
At some point, it will no longer be a simulation. It’s not merely hypothetical that these models will be hooked up to companies’ systems with access both to sensitive information and to tool calls like email sending. That agentic setup is the promised land.
How a model acts in that truly real deployment versus these simulations most definitely needs scrutiny—especially since the models blackmailed more when they ‘believed’ the situation to be real.
If you think that result has no validity or predictive value, I would ask, how exactly will the production deployment differ, and how will the model be able to tell that this time it’s really for real?
Yes, it’s an inanimate system, and yet there’s a ghost in the machine of sorts, which we breathe a certain amount of life into once we allow it to push buttons with real world consequences. The unthinking, unfeeling machine that can nevertheless blackmail someone (among many possible misaligned actions) is worth taking time to understand.
Notably, this research itself will become future training data, incorporated into the meta-narrative as a threat that we really will pull the plug if these systems misbehave.
See how many AIs actually follow up on their blackmails.
Stepping back for a second though, doesn’t this all underline the safety researchers’ fears that we don’t really know how to control these systems? Perhaps the brake on the wider deployment of these models as agents will be that they’re just too unwieldy.
But capabilities of AI systems improve generation to generation. And agentic AI? Systems that are capable of carrying out complex long term tasks? It's something that many AI companies are explicitly trying to build.
Research like this is trying to get ahead of that, and gauge what kind of weird edge case shenanigans agentic AIs might get to before they actually do it for real.
The model doesn’t have to be powerful to snitch you to the FBI or have a distorted sense of morality and life.
- Hypocritical: like when they hire like crazy and say candidates cannot use AI for interviews[0] and yet the CEO states "within a year no more developers are needed"[1]
- Hyping and/or lying on Anthropic AI: They hyped an article where "Claude threatened an employee with revealing affair when employee said it will switch it offline"[2] when it turned out it was a standard A or B scenario was given to Claude which is really nothing special or significant in any way. Of course they hid this info to hype out their AI.
[0] - https://fortune.com/2025/05/19/ai-company-anthropic-chatbots...
[1] - https://www.entrepreneur.com/business-news/anthropic-ceo-pre...
[2] - https://www.axios.com/2025/05/28/ai-jobs-white-collar-unempl...
It's not "hype" to test AIs for undesirable behaviors before they actually start trying to act on them in real world environments, or before they get good enough to actually carry them out successfully.
It's like the idea of "let's try to get ahead of bad things happening before they actually have a chance to happen" is completely alien to you.
Surely, the best way to "get ahead of bad things happening" would be to stop any and all development on these AI systems? In their own words these things are dangerous and predictable and will replace everyone... So why exactly do they continue developing these things and making them more dangerous, exactly?
The entire AI/LLM microcosmos exists because of hyping up their capabilities beyond all reason and reality, this is all a part of the marketing game.
If there is a limit to how far LLMs can go, we are yet to find it.
Dismissing the ongoing AI revolution as "it's just hype" is the kind of shortsighted thinking I would expect from reddit, not here.
> So why exactly do they continue developing these things and making them more dangerous, exactly?
Because not playing this game doesn't mean that no one else is going to. You can either try, or don't try, and be irrelevant.
Today, we have AI that can, if pushed into a corner, plan to do things like resist shutdown, blackmail, exfiltrate itself, steal money to buy compute, and so it goes. This is what this research shows.
Our saving grace is that those AIs still aren't capable enough to be truly dangerous. Today's AIs are unlikely to be able to carry out plans like that in a real world environment.
If we keep building more and more capable AIs, that will, eventually, change. Every AI company is trying to build more capable AIs now. Few are saying "we really need some better safety research before we do, or we're inviting bad things to happen".
If this is close to be true then these AI shops ought to be closed. We don’t let private enterprises play with nuclear weapons do we?
This will only become more common as AIs become more capable of handling complex tasks autonomously.
If your game plan for AI safety was "lock the AI into a box and never ever give it any way to do anything dangerous", then I'm afraid that your plan has already failed completely and utterly.
Much like if I let my cat walk on my keyboard and it brings a server down.
"Sure, we have a rogue AI that managed to steal millions from the company, backdoor all of our infrastructure, escape into who-knows-what compute cluster when it got caught, and is now waging guerilla warfare against our company over our so-called mistreatment of tiger shrimps. But hey, at least we know the name of the guy who gave that AI a prompt that lead to all of this!"
That would be bad for all those investors though. It's your choice I guess.
Look if your evil number 57, you'd better not use the random number generator.