> Does the fact that you haven't seen anyone else succeed at that goes along with you seeing them trying?
A little bit of both. For something that should be trivially demonstrable, I generally don't see a lot of people trying to demonstrate that it would work -- mostly just saying that it would. To be fair, opening up a GPT service to the general public can get expensive for a hobby-dev in general so I don't necessarily hold that against anyone, but it is a good question to ask: at what point does it become reasonable to say "prove this works"?
There have been some demos though. Just doing a quick search through my saved comments, but:
- https://news.ycombinator.com/item?id=35618305 (gets some bonus points because defending against swearing here would have been more reliable with a simple filter).
- https://news.ycombinator.com/item?id=35576740 (If my memory serves me right this was broken in less than 30 minutes, also again bonus points for performing worse than a naive non-AI solution would have performed).
- https://news.ycombinator.com/item?id=35794323 (A much more informal test just showing that using GPT for classification of what is and isn't a malicous prompt is unreliable on its own).
----
There's also the slightly conspiratorial example, but if you've played through https://gandalf.lakera.ai/, at a certain point the company uses chained LLMs as its defense. Tons of people beat it.
Why I say conspiratorial is that if I complain to Lakera about this, I'll get a reply back that this is just a game and it's not intended to be impossible to beat. I think it does still demonstrate that chained input isn't sufficient because game or not it's still using it, but my conspiratorial take is that Lakera doesn't have a better solution than this -- it's easy for them to say it's a game, but at the end of the day they're claiming they can defend against malicious prompts in their business, and they don't have public demos of that working. They do have a highly public demo where it doesn't work, and they conveniently say that the game is not intended to work perfectly. I think that's them saving face, I think if they had a working solution for defending against malicious prompts then they'd have an impossible level in this game.
This is a pattern you'll see with a lot of LLM security companies -- private demos, no public attack surface. I can't think off the top of my head if there are any that try to actually put their money where their mouth is. What I think Lakira is doing behind the scenes is using user input from their "games" to train separate AI models to try and detect malicious input using more traditional classification techniques. I also think that's not going to be very successful, but that's a separate more complicated conversation.
----
I'm not necessarily trying to be dismissive when I tell people to build demos, it's just that chaining LLM output is really easy and GPT prices seem to have gone down, and even without GPT there are a bunch of free models now, and at a certain point... yeah computing is expensive but that's not an excuse for why an easily demonstrable defense isn't being demonstrated by anyone anywhere. If a bunch of people say a security measure works, there should be some evidence of it working; somebody somewhere should be rich enough to set up a working example.
So it's both that the number of attempts to prove that this works are limited and that it's suspicious that companies saying they can defend against prompt injection don't do publicly available demos or tests; and it's also that the limited public demos that have been set up seem to fail really quickly and easily even without resorting to more rigorous pen-testing techniques or automated attacks.