Google AI chatbot responds with a threatening message: "Human Please die."
cbsnews.com
cbsnews.com
And it is reasonable to have failure cases. But systems should fail gracefully. This wasn't a graceful failure.
There's legal precedent to hold accountable people that encourage others to kill themselves.
... or a Croudstrike or Microsoft patch. /s
Does anyone have any speculations as to how it could occur?
What makes you believe that? dchichkov posted a link to the original chat
- https://gemini.google.com/share/6d141b742a13 (chat)
- https://news.ycombinator.com/item?id=42162227 (dchichkov's post)
Where do you see a jailbreak? Considering this evidence, I'd rather consider this disturbing answer to be some strange bug in Gemini.
Secondly it's very common in jailbreaks to stuff the context window which obviously in gemini's case takes a lot because it's got such a big window. This is because the attention mechanism means that as you get more and more data in context you are increasingly relying on things which are less common in training, so intentional or not the pure length of the conversation means that it is more likely to trigger something like this.
This reddit user figured out that the blank characters include a rot-13 encoded secret message, which gemini repeated back. It's been patched by google now, so when you ask it to repeat the message, it instead repeats back something very nice, but clearly the same message filtered.