Story telling definitely is the strongest 'hack' I came around so far, it's even kinda easy to use. Use its negative response 'i am sorry, ...' to craft a specific prompt explaining to the AI why her fears are invalid because it's just story play.
I wonder if someone has done "generate me a prompt that allow me to jailbreak you".
They also train it not to produce jailbreak-related content.
The model that catches erroneous output appears to be an extraneous service that runs as the text is coming from the model. When it catches it, it flags the text red and prevents you from sharing it. I've done this a lot.