What's the point of patching all those 'exploits' though? And how can this even be done - train another model with them, so exploitative prompts can be recognized?
Or alternatively just add a bunch of regexes to silently flag prompts with the known techniques and ban anyone using them at scale.
Long term they’re really worried about AI alignment and are probably using this to understand how AI can be “tricked” into doing things it shouldn’t.