I wrote about why I don't think that will work here: https://simonwillison.net/2022/Sep/17/prompt-injection-more-...
The problem with this approach is that prompt injection is an adversarial attack.
A statistical approach that catches 99% of possible attacks is worthless, because a bunch of people on a subreddit somewhere will keep on plugging away at it until they find a hole - and will then share the hole they've found like wildfire.
This isn't a theoretical problem: it's happening already. Look at how the whole DAN thing came together: https://kotaku.com/chatgpt-ai-openai-dan-censorship-chatbot-...
If you showed me a SQL injection mitigation attack that only worked 99% of the time I would laugh at how naive you were being!