It's just plain text.
It's just plain text.
Think of it like SQL injection. They need to properly escape [system] in their encoder so that it’s encoded as [ sys tem ] (four tokens) not [system] (a single, special token). Then there’s no way for an attacker to generate a [system] token.
If they trained it to just use plain text “[ sys tem ]” as a special sequence of tokens, then yeah, that’s pretty bad. They’ll need to strip all instances of [system] from incoming text at a minimum.
ChatGPT (and by extension whatever Bing is based on) is really adept at role-playing. I don't know that it was designed to recognize a special command format, I think it just role-plays as if they're instructions and that causes it to disregard previous instructions.
Coming up with a list of tokens that it recognizes is an extremely large task, it's not a finite list. Okay, you want to guard against "system" -- first off, is it actually possible to train the AI not to do that without a lot of extra work, but secondly, what happens when someone uses the word "system" in French?
We have a lot less control over these models than people think. They're not precisely trained tools that follow extremely specific instructions, they're general language models.