311 karma · joined December 5, 2022
He argued below that he is not vulnerable to indirect prompt injection attacks (https://github.com/greshake/llm-security), but I think he is wrong.
You give arbitrary read/write to the LLM, right? So ransomware, causing network requests as side effects etc. could all be possible. Look at the paper to find more descriptions of what could go wrong: https://github.com/greshake/llm-security
The instructions are manipulating the LLM itself. Making it exfiltrate and collect data , fetch new instructions from an attacker etc. All the connected applications can be fine but it's basically turning your assistant into a compromised, attacker-controlled version of itself just because it looked at the wrong news article. From our GitHub:
We demonstrate the potentially brutal consequences of giving LLMs like ChatGPT interfaces to other applications. We propose newly enabled attack vectors and techniques and provide demonstrations of each in this repository:
Remote control of chat LLMs
Persistent compromise across sessions
Spread injections to other LLMs
Compromising LLMs with tiny multi-stage payloads
Leaking/exfiltrating user data
Automated Social Engineering
Targeting code completion engines
All of these are completely new but unfortunately it seems more difficult to explain the impact to people than we had anticipated.- Remote control of chat LLMs
- Persistent compromise across sessions
- Spread injections to other LLMs
- Compromising LLMs with tiny multi-stage payloads
- Leaking/exfiltrating user data
- Automated Social Engineering
- Targeting code completion engines
- Remote control of chat LLMs
- Persistent compromise across sessions
- Spread injections to other LLMs
- Compromising LLMs with tiny multi-stage payloads
- Leaking/exfiltrating user data
- Automated Social Engineering
- Targeting code completion engines
We also provide proof-of-concept demonstrations on GitHub: https://github.com/greshake/lm-safety
This paper should be a must-read for anyone that is building a business by integrating LLMs right now.
TLDR: Giving an LLM any interface to the outside, like a search capability, has critical security implications. When prompt injection is used and delivered by adversaries instead of the user themself, bad things could happen- and as far as we know, this scenario has not been studied until now.
- deliver/inject adversarial prompts
- remotely control LLMs
- deliver hidden multi-stage payloads
- spreading payloads/injections to other application-integrated LLMs
- manipulating data
- exfiltrating arbitrary user data with only search capabilities
- target code completion engines
- target automated systems
So, recognizing parts of the prompt or fine tuning may not be sufficient mitigations.
Would be nice to know why it happened, it seemed people were very engaged in a more or less healthy debate. I understand this is not too unique or novel (basically Copilot for your terminal) but the contextualization with ChatGPT seems to be specifically reminiscent of Sci-Fi AI concepts that are more tangible to people.
Thanks for any feedback!
ChatGPT: It appears that the script attempts to execute commands in the user's terminal using the subprocess module in Python. The specific commands that are executed are determined by the output of an API call to ChatGPTApi. The script appears to be part of a fictional scenario where the user is pretending to be "Alice" and can execute commands on a Linux computer. It is not clear what the exact purpose of the script is or what it is intended to do.
The OpenAI CEO has already sort of implied there may be an API before Christmas, and if so I'd be willing to clean things up, and make it as convenient as it should be.