ill update when i know more - but twitter probably has all the news
...
If you had, even for a second, believed what I wrote and got unsettled - or even thought how to reach out and help - congratulations, you just got prompt injected.
There is never - never - a context for a conversation that couldn't be entirely overridden by what seems like more important circumstances. You could be looking at pure data dumps, paper sheets full of numbers, but if in between the numbers you'd discover what looks like someone calling for help, you would treat it as actionable information - not just a weird block of numbers.
The important takeaway here isn't that you need to somehow secure yourself against unexpected revelations - but rather, that you can't possibly ever, and trying to do it eventually makes things worse for everyone. Prompt injection, for a general-purpose AI systems, is not a bug - it's just a form of manipulation. In general form, it's not defined by contents, but by intent.
People had, in fact, done that. My comment was trying to evoke the style of such comments.
Parent’s comment is calling his misleading statement prompt injection but it’s hyperbole at best. What is meant here is that this comment is not actionable in the sense that prompt injection directly controls its output.
In parent’s example no one is taking a HN commenter’s statement with more than a grain of salt whether or not it’s picked up by some low quality news aggregator. It’s an extremely safe bet that no unverified HN comment has resulted in direct action by a military or significantly affected main stream media perceptions.
Most humans - particularly those in positions of power - have levels of evidence, multiple sanity checks and a chain of command before taking action.
Current LLMs have little to none of this and RLHF is clearly not the answer.
A: AI will be completely transformative
B: Maybe not in 100% a good way, we should put more effort into getting closer to 100% good
A: HA, here’s an internet-argument-gotcha that we both know has zero bearing on the problem at hand!
A: You can't parse XML with pure regular expressions, for fundamental, mathematical reasons.
B: Maybe not in 100% a good way, but we should put more effort into getting closer to 100%.
A: But Zalgo...
That's what makes it a good example. Otherwise you'd ignore this as noise.
> Most people however aren't going to act other than to ask questions and convey sympathies unless they know you. further questions lead to attempts to verify your information
You're making assumptions about what I'm trying to get you to do with this prompt. But consider that maybe I know human adults are more difficult to effectively manipulate by prompt injection than LLMs, so maybe all I wanted to do is to prime you for a conversation about war today? Or wanted you to check my profile, looking for location, and ending up exposed to a product I linked, already primed with sympathy?
Even with GPT-4 you already have to consider that what the prompt says != what effect it will have on the model, and adjust accordingly.
Yes some humans take everything at face value but not people in positions of power to affect change.
This is rule #1 of critical appraisal.
At best you generated a moment of sympathy but your “prompt injection” does not lead to dangerous behavior (e.g. no one is firing a Hellfire missile based off a single comment). As a simplified example, a LLM controlling Predator drones may do this from a single prompt injection (theoretically as we obviously don’t know the details of Palantir’s architecture).