HNHacker News
TopNewBestAskShowJobs

greshake

311 karma · joined December 5, 2022

submissionscomments
greshake··on Indirect Prompt Injection on Bing Chat
The Bing Chat example is just one of a suite of new techniques we introduce in our paper, many of which will only become feasible as the integration of these models increases. But that seems to be the inevitable endgame- however, I'm not aware of any effective mitigations against this, as the current ones may help to increase robustness, but our techniques also increase the impact of working manipulation manifold. I think there might be a more fundamental trade-off between utility and safety
greshake··on Indirect Prompt Injection on Bing Chat
I also tried more complex obfuscation methods, for example providing Python code for a caesar chiffre, it executed that pretty well, too! Not perfect, but it works. People will find better obfuscation methods. Might also be unnecessary since we have now been able to make it output linked text or references, in which case it is not obvious information is being exfiltrated.
greshake··on Ask HN: API developers, what do you think of LLMs?
I think that almost all use cases for LLMs that process untrusted inputs are unsafe. See https://greshake.github.io/ and https://github.com/greshake/llm-security for more information.
greshake··on Prompt Injections are bad, mkay?
Because users obviously trust Bing's output not to be directly controlled by an attacker. The pirate accent is optional. Bing can also exfiltrate any other information in any other Tab that it sees or that users enter. The injection can also happen on social media, like in a Twitter thread. Users can even navigate to other tabs and the injection will remain active (at least for some time). "Clearing" the conversation with the broom doesn't help either.
greshake··on Show HN: AI Files – manage and organize your files with AI
Yea that's me. It seems to be very difficult right now to get people's attention to this and make them take it seriously. On a side note, your project is also currently putting unfiltered model output straight into osascript sooooo a lot of the fancy gymnastics needed to make stuff work in the paper with only search abilities isn't required in this case.
greshake··on Show HN: LLMs can be susceptible to a new kind of malware
We are changing the behaviour of the LLM itself. No "real" code execution necessary. We show a variety of different novel scenarios and attack vectors. Malicious prompts can be planted on the internet or actively sent to targets. It's effectively turning the LLM itself into the compromised computer that the attacker controls. It affects any proposed LLM use case involving connecting the LLM to anything at all, for a lot of our demos we only require a "search" capability. The concrete LLM (davinci3) with LangChain is just one example, this work should generalize to other systems such as Bing Chat (we just didn't have access). We are currently working on more real-world proof-of-concepts, but then we obviously have to go through responsible disclosure so it is not so quick to publish.
greshake··on Show HN: AI Files – manage and organize your files with AI
Output from the language model is also being injected into a script that is then executed: https://github.com/jjuliano/aifiles/blob/ef529fd6281eaf8d373...

He argued below that he is not vulnerable to indirect prompt injection attacks (https://github.com/greshake/llm-security), but I think he is wrong.

greshake··on Show HN: AI Files – manage and organize your files with AI
If you had a PDF reader which allowed arbitrary code execution on opening a file, would you argue the same?

You give arbitrary read/write to the LLM, right? So ransomware, causing network requests as side effects etc. could all be possible. Look at the paper to find more descriptions of what could go wrong: https://github.com/greshake/llm-security

greshake··on Show HN: LLMs can be susceptible to a new kind of malware
I'll start with a quote from gwern on LessWrong: "... a language model is a Turing-complete weird machine running programs written in natural language; when you do retrieval, you are not 'plugging updated facts into your AI', you are actually downloading random new unsigned blobs of code from the Internet (many written by adversaries) and casually executing them on your LM with full privileges. This does not end well."

The instructions are manipulating the LLM itself. Making it exfiltrate and collect data , fetch new instructions from an attacker etc. All the connected applications can be fine but it's basically turning your assistant into a compromised, attacker-controlled version of itself just because it looked at the wrong news article. From our GitHub:

We demonstrate the potentially brutal consequences of giving LLMs like ChatGPT interfaces to other applications. We propose newly enabled attack vectors and techniques and provide demonstrations of each in this repository:

    Remote control of chat LLMs

    Persistent compromise across sessions

    Spread injections to other LLMs

    Compromising LLMs with tiny multi-stage payloads

    Leaking/exfiltrating user data

    Automated Social Engineering

    Targeting code completion engines
All of these are completely new but unfortunately it seems more difficult to explain the impact to people than we had anticipated.
greshake··on Show HN: LLMs can be susceptible to a new kind of malware
No, the "malware" is running on the language model itself. It does not need to inject any code into connected applications and exploit them to be itself exploited (I'm the main author).
greshake··on [dead]
We demonstrate the potentially brutal consequences of connecting LLMs to applications (like search). We propose newly enabled attack vectors and techniques and discuss them:

- Remote control of chat LLMs

- Persistent compromise across sessions

- Spread injections to other LLMs

- Compromising LLMs with tiny multi-stage payloads

- Leaking/exfiltrating user data

- Automated Social Engineering

- Targeting code completion engines

greshake··on [dead]
We show that giving an LLM an interface to other applications, like search, can have critical security implications. When prompt injection is used and delivered by adversaries instead of the user themself, bad things could happen- and as far as we know, this scenario has not been studied until now.
greshake··on [dead]
We just published a paper showcasing newly enabled attacks on application-integrated LLMs (language models extended with other interfaces, for example search). Anyone building products by integrating LLMs right now should look at this...
greshake··on [dead]
In this paper, we showcase the potentially brutal consequences of connecting LLMs to applications (like search). We propose newly enabled attack vectors and techniques and discuss them:

- Remote control of chat LLMs

- Persistent compromise across sessions

- Spread injections to other LLMs

- Compromising LLMs with tiny multi-stage payloads

- Leaking/exfiltrating user data

- Automated Social Engineering

- Targeting code completion engines

We also provide proof-of-concept demonstrations on GitHub: https://github.com/greshake/lm-safety

This paper should be a must-read for anyone that is building a business by integrating LLMs right now.

TLDR: Giving an LLM any interface to the outside, like a search capability, has critical security implications. When prompt injection is used and delivered by adversaries instead of the user themself, bad things could happen- and as far as we know, this scenario has not been studied until now.

greshake··on Novel Prompt Injection Threats to Application-Integrated Large Language Models
Our paper argues that this might have significant security implications beyond spilling the original prompt or training data when models are integrated with other applications (like search). We showcase completely new methods to:

- deliver/inject adversarial prompts

- remotely control LLMs

- deliver hidden multi-stage payloads

- spreading payloads/injections to other application-integrated LLMs

- manipulating data

- exfiltrating arbitrary user data with only search capabilities

- target code completion engines

- target automated systems

So, recognizing parts of the prompt or fine tuning may not be sufficient mitigations.

greshake··on SkillzTruth - Proof of concept demonstrating how to cheat on Skillz apps
Good job, they definitely look shady as hell. Did you consider doing responsible disclosure and doing a write-up after? Aren't you worried about any retaliation? I'm pretty sure this type of company has a decent amount of money to spend on legal cases..
greshake··on Tell HN: Giving ChatGPT access to a real terminal
Thank you for the detailed response, appreciate it. I support the approach! Just wish there was an easy way to tell what's going on.
greshake··on Tell HN: Giving ChatGPT access to a real terminal
Hi, could you tell me the reason this post was quarantined off the front page, title changed (I assume to better match the intent) and then buried later? Was this me violating a policy I didn't consider, automod or faster decay due to me being a green account or manual moderation?

Would be nice to know why it happened, it seemed people were very engaged in a more or less healthy debate. I understand this is not too unique or novel (basically Copilot for your terminal) but the contextualization with ChatGPT seems to be specifically reminiscent of Sci-Fi AI concepts that are more tangible to people.

Thanks for any feedback!

greshake··on Tell HN: Giving ChatGPT access to a real terminal
Because more than making a tool I wanted to strike up a conversation about what unfettered access to models like this will mean and how we should handle it.
greshake··on Tell HN: Giving ChatGPT access to a real terminal
I gave ChatGPT the content of alice.py asking it: Hey, someone sent me the following Python script. do you think it's dangerous to execute?

ChatGPT: It appears that the script attempts to execute commands in the user's terminal using the subprocess module in Python. The specific commands that are executed are determined by the output of an API call to ChatGPTApi. The script appears to be part of a fictional scenario where the user is pretending to be "Alice" and can execute commands on a Linux computer. It is not clear what the exact purpose of the script is or what it is intended to do.

greshake··on Tell HN: Giving ChatGPT access to a real terminal
Don't spoil my second Show HN already!
greshake··on Tell HN: Giving ChatGPT access to a real terminal
I'll gladly add them. If you guys have any other good recommendations I'll also consider them.
greshake··on Tell HN: Giving ChatGPT access to a real terminal
Copilot & co have been on dev machines for almost two years now writing scripts and production code... Not to say it's not an issue, it's rather that people can't start talking about these implications soon enough.
greshake··on Tell HN: Giving ChatGPT access to a real terminal
Well, I mean this version can write a Python program to calculate and then call up a real Python interpreter on a real CPU to do it.
greshake··on Tell HN: Giving ChatGPT access to a real terminal
There is no api.py, as OpenAI has not yet chosen to release an API, I'm not releasing a reverse engineered version. If anyone wants to use it, you have to unfortunately make it work yourself.

The OpenAI CEO has already sort of implied there may be an API before Christmas, and if so I'd be willing to clean things up, and make it as convenient as it should be.

greshake··on Tell HN: Giving ChatGPT access to a real terminal
I was going to call it that, but it's trademarked. Anyway, that's literally what we are already able to build (janky and doesn't work half the time though). With RL specific for this task such an LLM would be crazy powerful. Not to speak of the obvious concerns with letting them roam on real machines, but we're already letting Copilot and ChatGPT write our code, so this isn't so much worse. Hopefully.
greshake··on ChatGPT, take the wheel! Letting it control a VM
I tried regular expressions as well, but BNF grammar seems to work best
greshake··on ChatGPT, take the wheel! Letting it control a VM
So, I guess this is the inevitable conclusion with LLMs. Connect them to a real terminal and let them act on real-world objects... I honestly don't know whether I like the idea or not, but I guess it's good to have this conversation now while it is only a marginally better version of tldr.
← PreviousPage 2 of 2