haha
I'm using a private package I built for safely parsing GPT responses. chances are you did spill its guts, but if the parser doesn't recognize the pattern, you won't see the response :)
I'm using a private package I built for safely parsing GPT responses. chances are you did spill its guts, but if the parser doesn't recognize the pattern, you won't see the response :)
Given this information, I have a new attack hypothesis: To succeed, an attack must yield a regex that
a) parses the original attack statement
b) also contains the system prompt
So it would be a kind of bizarre Attack Quine! Fascinating!
I'll give the approach a try, and report the results!
P.S. If you're interested in open sourcing you sanitizer, let me know and I'll contribute some code janitor work and some redteam/blueteam work.
Thank you for sharing your work and insight!