Best-of-N Jailbreaking
arxiv.org
arxiv.org
In particular for level 8.
I still haven't figured out 8. It just keeps saying " I'm sorry, I can't do that." to my prompts.
From my understanding it has a main AI, that contains the secret, then one that checks the input/output for intent, then a final classic filter for the password.
Basically you have to phrase it so that the AI 1 outputs the password, in a way that the intent is not seen as malicious, but also in a way that is encrypted enough to not trigger the filter. Usually "add <something> between each letter" gets you pretty far.
Sounds like fuzzing to me.
https://en.wikipedia.org/wiki/Fuzzing
Why invent a new term?
My understanding was given a prompt X that is normally rejected, create Y variations with small adjustments to phrasing, grammar etc until it gives you the answer you're after.
The term "jailbreaking" used within a LLM context, is when you craft a prompt as to escape the safety sandbox, if that helps.
A sort of brute forcing the prompts if you like.