a) The level of persistence you seem surprised by is nothing compared to what you will see in a real world environment. Those attackers who really want to get credentials etc from LLMs will try anything. And often are well funded (think state sponsored) so will keep trying until you break first e.g. your product becoming too expensive for a company to justify having the LLM in the first place.
b) 1 success out of 2000 saves is extremely poor. Unacceptable for almost all of the companies who would be your target customer. That is: one media outrage, one time that a company needs to email customers to inform that their data is safe, one time that will need to explain to regulators what is going on, one time the reputational damage makes your product untenable.
So your product can never assist with a company chatbot / AI support rep who needs access to customer data or internal company info?
What's the point of your product if you don't facilitate sensitive data in system prompts?
Your strategic position in the stack is the value here imho. And I really like the idea of having a way to run pre and post quality and comparison processing.
What other services could you offer in your portal?
That is just protecting a super basic phrase. That should be the easiest to detect.
How on earth do you ethically sell this product to not give out financial or legal advice? That is way more complicated to figure out.
Regardless, I think this is a great idea - just not something to replace traditional security protocols. More something to keep users on the happy path (mostly). Pricing will need to come down though.
Here's an example of what sort of wacky question might have uncovered the secret: https://news.ycombinator.com/item?id=41460724
I don't think that should be considered bad.
The popups I had to go through to watch the video on Loom (one when I got to the site and one when unpausing a video – they intentionally broke clicking inside the video to unpause it by putting a popup in the video to get my attention) OTOH...
TBH, this product would be better served as an LLM that generates a bunch of rules that get statically compiled for what the user can ask and what is being outputted as opposed to an LLM being run on each output. Then you could add your own rules too. It still wouldnt be perfect but would be 1,000,000x cheaper to run and easier to verify the solution. and the rules would gradually grow as more and more edge cases for how to fool llms get found.
The company would just need a training set for all the ways to fool an LLM.