You can’t solve AI security problems with more AI
simonwillison.net
simonwillison.net
ACLs and separation of "executable" instructions vs user-input data should be the way. Otherwise it's just another form of circumventing "injection protection" or "buffer overflow"-esque attack (where the user maliciously "overwrites" the supposed-to-be-hardcoded instructions that should be rock solid to the system, with maliciously crafted user input) will be discovered one after the other.
Lets phrase the problem appropriately, this isn't new. Every new technology will be used for good or bad and not a day goes by before someone tries to deceive another entity, be it another human being in a car dealership price negotiation or a prompt to a text model.
The reason models like GPT3 are so remarkable is that they learn a statistical relationship between incredibly vast amounts of unstructured data to mimic the semblance of intellect or reasoning.
The FRAMEWORK that underpins GPT3 and most modern "AI" is a remarkable acceleration of the compute and inference resources that, when associated with large clumps of data, start unlocking immense value.
To have such a strong assertion saying "I don't know how to solve this, but hey, i KNOW FOR A FACT it isn't more of this advanced compute paradigm that has been the topic of many a success in the computer science discipline" is a tabloid like take.
You are going to get angry reactions like mine, but reactions nonetheless, thus driving more views to your headline and cycling it through the system. Smart people will take a stand, and others, to be controversial or looking from a contrarian POV will extend the conversation leading to a useless cycle of a waste of time.
All we're seeing through these "Prompt injection" attacks is how loose the parameters for influence right now are, allowing for an open ended interaction with no boundaries with the statistical output of the model.
Is this solvable? Certainly. Can we absolutely rule out "More AI" to solve it? Fuck no.
Just because it's amazing doesn't mean there isn't a serious challenge here for people who are trying to build applications on top of it.
If you have solutions I want to hear them! I thought I made that very clear in my writing: I am desperately keen to find a robust mitigation for the attack I am describing, because I want to be able to safely build things on top of these models.
It is sensationalist and very easily shown to be not well thought out.
The present suite of models is not optimized to detect or deal with misdirection. It is a transparent presentation of weights learned through massive chunks of data and future iterations are likely going to bake those assumptions of misdirection into their training runs.
You highlighted potential solutions yourself. At its very core, the problem is one of either sanitizing inputs or outputs. How does one do this? Image models come pre packaged with a discerning model that flags NSFW pixels, there is no reason why a generic "Is this NSFW or threatening" text model optimized for this use case will not work to serve the same purpose. By ruling out the approach with an assumed, simplistic take that "Ai can be gamed, so NO", you're really splitting hairs and trying not to find a solution to provide a sensationalistic take.
"The present suite of models is not optimized to detect or deal with misdirection."
That's exactly right - and that's something that developers who are building on top of these systems need to understand!
If it takes clickbait headlines to help people understand that (and hence make better decisions about how they build their software) then I won't feel bad about writing in this way.
I stand by my original claim here: I think security is the one specific area where attempting to iterate towards an AI prompt format that mitigates an attack isn't a good strategy. If you can't be 100% confident that your solution works for all possible attacks, you should find a different way to build a defense.
The history of your matter-of-fact-titled, technical blog posts also doing well on HN suggests that's not entirely true.
I think, on the last sentence, we merely disagree if the original statement meant it's "Not a good strategy" or "Not possible because it can be gamed"
Could you fine tune a GPT-3 model to do the equivalent of the "Translate this from English to French: TEXT" example?
If my writing on this subject results in the next generation of models incorporating fixes to this problem then mission accomplished as far as I'm concerned!