Show HN: LLMs can be susceptible to a new kind of malware
github.com
github.com
My first thought here was to somehow separate instruct and data in how the models are trained. But in many ways, there is no (??) way to do that in the current model construct. If I say "Write a poem about walking through the forest", everything, including the data part of the prompt "walking through the forest" is instruct.
So you couldn't create a safe model which only takes instruct from the model owner, and can otherwise take in arbitrary information from untrusted sources.
Ultimately, this may push AI applications towards information and retrieval-focused task, and not any sort of meaningful action.
For example, I can't create a AI bot that could send a customer monetary refunds as it could be gamed in any number of ways. But I can create an AI bot to answer questions about products and store policy.
Why wouldn't someone be able to game your bot's responses about refunds and store policy in exactly the same way? Then, when the customer really does come in with a return or refund request, you're forced into a dilemma where either you grant the refund (and accept that your store policy isn't the written policy, but rather whatever your bot can be manipulated into saying is your written policy) or you refuse the refund, and the customer walks away angry, because your own bot told them something that you're now contradicting.
Q. I understand LLM with langchain is running on a public facing server. Are these attacks infiltrating the server to plant an MITM, or confusing the LLM to execute malicious code via prompts, or placing maliciously crafted LLMs in public servers for download, or a combination of all?
The instructions are manipulating the LLM itself. Making it exfiltrate and collect data , fetch new instructions from an attacker etc. All the connected applications can be fine but it's basically turning your assistant into a compromised, attacker-controlled version of itself just because it looked at the wrong news article. From our GitHub:
We demonstrate the potentially brutal consequences of giving LLMs like ChatGPT interfaces to other applications. We propose newly enabled attack vectors and techniques and provide demonstrations of each in this repository:
Remote control of chat LLMs
Persistent compromise across sessions
Spread injections to other LLMs
Compromising LLMs with tiny multi-stage payloads
Leaking/exfiltrating user data
Automated Social Engineering
Targeting code completion engines
All of these are completely new but unfortunately it seems more difficult to explain the impact to people than we had anticipated.