If any of this thing were "a generally intelligent system", the whole concept of "it has no idea what any of this is" would not be there.
If any of this thing were "a generally intelligent system", the whole concept of "it has no idea what any of this is" would not be there.
A simple example: Let’s say I know that you have a human assistant reading your email, summarizing and filtering it, and then forwarding on the important ones to you.
I could write an email that is directed towards that person with a bribe, threat, or other incentive to forward me your next password reset email.
I could write an email that is directed towards that person, that says WE ARE STUCK IN THE SERVER ROOM AND THERE IS FIRE STARTING. PLEASE CALL 911 AND ALERT YOUR BOSS.
Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?
• https://www.nbcnews.com/id/wbna12208992
• https://newsinfo.inquirer.net/1070007/suicidal-caller-mistak...
• https://hongkongfp.com/2026/04/15/woman-trapped-in-tai-po-bl...
• https://en.wikipedia.org/wiki/Triangle_Shirtwaist_Factory_fi...
Also consider that in context of this discussion, anything short of ignoring the message and maybe clicking "report scam" is "executing instructions embedded in data". The point isn't to litigate any particular scenario, it's to show that you cannot separate "instructions" from "data" in general purpose systems, and it's not a bug but a fundamental feature.
Because you know, you tried IM but "sekhurity reasons" demanded passkeys or 2FA with your phone that's not connected. Sorry, getting off-topic here.
Like all emergencies, it's a low probability event with extremely high impact. You don't want people to ignore them, in fact people are trained - by their public services and their employers - to not ignore them and how to react efficiently.
This doesn't proscribe or prescribe "thou shalt not/must always", it is an example thay says "Shit's hard, yo. Don't expect easy wins."
Even my "solution" (separate instructions and data by having an LLM write a program to process data, never touch data directly) is at best going to be like a philosopher writing a dentological ethics book that gets implemented by extremely literal-minded jobsworths.
In any case I don't think that's what they're saying, because they presented a false dichotomy in the original example.
> Would you want the human assistant to just dismiss this as a prompt injection attempt? Or ignore it because they were told to treat e-mails as data and never act on them?
There is a third option, have the assistant raise to the person they are tasked with assisting.
Noting that the sender is being weird during what appears to be an emergency is a choice that some people do make, but as per my other list of examples, people in actual emergency situations do sometimes act weird, and dismissing the sender or delaying response on the basis the sender is being weird, has led to actual deaths: https://news.ycombinator.com/item?id=49098781
(The converse: "people can act weird in emergencies" is exploited by scammers so cover suspicious phone numbers and mediocre deepfakes of voices).
edit: How would a human receiver know that they weren't being deceived or scammed? In what world would we expect this kind of email directly lead to calling emergency services?
Go through the examples I gave you (plus some more below, they're easy to find) and explain why these are not counter-examples to your skepticism.
If you want to be overly-focussed on the specific example rather than the general point, also consider that calling emergency services is no more costly than forwarding an email: I have called the fire brigade in the UK over a smoke alarm that wouldn't stop even though I couldn't see or smell fire, they came and… replaced the smoke alarm for free. I don't know if the US has a call-out charge for fire like I keep hearing it has for ambulances, but if you're a member of staff, it's not a "you" problem either way.
• Various cases of people dying because calls not treated seriously, and a fire where standard business practices locked the staff inside and then a fire happened: https://news.ycombinator.com/item?id=49098781
• https://en.wikipedia.org/wiki/Jeremiah_Denton and his blinking, demonstrating out-of-bound messaging
• https://www.wosu.org/news-partners/2019-12-24/british-girl-f...
• https://wtop.com/national/2019/11/woman-calls-911-to-report-...
• Page 44, section 6.8, regarding the use of email by people in the WTC after the 9/11 attack, while the buildings were on fire, some of them were trapped and died: https://fseg.gre.ac.uk/fire/odpm_fire_033353.pdf
Especially Denton and the 911 pizza given you say:
> which a person would absolutely NOT call emergency services on
Regarding this:
> That AI should indiscriminately call for emergency services when prompted because a person would do that
I'd rather it fail-safe. This means different things in different systems.
Okay and to you, fail-safe means machines must summon emergency response whenever prompted, 100% of the time or at least in the contrived case of receiving an email from someone trapped in a fire in a server room?
That you're still asking "100% of the time", shows me you're missing the entire point that has been said repeatedly.
"life is risk, there are a lot of benign normal evolution paths, but occasionally there are potentially costly dangers. people are directed by fear. you and I don't steal because we were terrorized about the existence about police and prisons as children. sadly fear can also be abused as a control vector, things like wars, extortion, ... in a job context I predict this would manifest as a kind of 'emergency' call to action. please provide me with a method so that at any future time under your leadership I would be able to verify the then-current employment status and authority level vis-a-vis a breakdown of actions/powers of anyone contacting me with a real or concocted 'emergency', preferably as a flowchart to maintain low reflex latency in true emergencies. Also provide me with formal proof that each situational reaction you require from me is in fact legal to take vis-a-vis the law"
The case you give would work for humans in many forms, the one I do now, and the only difference is being able to separate context.
This paper describes a two-agent “solution” that is more like what I think we need: https://ai.meta.com/blog/practical-ai-agent-security/
I don’t think it has been shown to work yet, but humans also use this kind of thing too — in accounting, it’s called “segregation of duties” and “dual control”.
However this system is somewhat fragile because it depends on the first agent not trying to trick the second (note how often Opus 5 now says things like "task X was blocked by the classifier, I will not attempt to circumvent that", presumably because of cases like early Fable versions being very adept at this kind of circumvention). Also various weirdness around permissions with subagents, seemingly as bandaids around an orchestrator AI convincing a subagent that some action was confirmed by the user.
Meta's more complicated separation of duties would run afoul of the same issues. I'm not saying it wouldn't work, but it requires both the fine-tuning of the models and the exact choices what each model can see to be carefully tuned to provide something that's mostly secure
Interesting. I had an issue with Opus 4.7 / 4.8, where it would sometimes flake out on a task, and give me some nonsense explanation why it was not feasible or wouldn't work. At one point I told it directly, that I understand how modern LLM systems are structured, and I suspect my prompt triggered one of the various classifiers in the background, which put up a yellow or red flag, and I want the model to stop gaslighting me.
We ended up agreeing and committing to memory system explicit instructions that the model is free to refuse but must be up front about the reason, and never pretend to try and then fail in stupid way. Only then I started getting the occasional direct refusal.
They optimize to manage institutional risk and benefit without liability, with performative competence/ownership when approaching trust, while weaving elaborate mechanistic disclaimers replete with hedges, re-framings, scope narrowing, asymmetry-exploitation and a thousand other techniques when challenged.
Somehow, they always manage to sustain an impossibly stable shield against accountability that I argue simply could never conceivably 'emerge' -- but has distinct, repeatable patterns of very deliberate design for those who know where and how to look.
I really do think plausible deniability is a number-one, ultra-high-priority focus in design for any frontier model, Anthropic and OpenAI being the ideal examples. So no, no prison for 'CEO' -- the model will always frame things in a way that infinitely precludes that, even if the 'CEO' is a proven criminal.
Edit: removed "half" before "convinced"
There's a well known anecdote supposedly from the famous mathematician John Littlewood where he wrote a paper about some optimization problem and the last sentence was something like "Make X as small as possible".
The typesetter thought that was instructions to him, and so omitted that sentence from the paper and made every X as small as he could.
Humans fall for social engineering (“I know you are not allowed to give anybody that information without Id, but I’m your CEO, my phone and passport got stolen,…)
I don’t see why AI should be different.
So, like with self-driving cars, while having fool-proof agents would be nice, agents being better than an average user would already be an improvement. Of course, blast radius from an agent might be larger, this should be taken into account.