This might fool an instruction tuned LLM. But not a lowly T5.
I agree that you won’t catch 100 %. But you also spoke about how having these silly rules in your prompt against leaking and then making it easy for your users to fool the model into leaking that very rule so they can post it on their Twitter is embarrassing.
Using a pre-filter that is not LLM-based (and maybe even counting the number of injection attempts, deliberately outputting fake prompts, etc., to really muddy the water for anyone trying) - that’s just the kind of nod ti show “Hey hacker guys, we’re not noobs here”. Kinda like the companies that put hiring messages into their website’s source code. Not about protection, really. But respectability.