With plugins, GPT-4 posts GitHub issue without being instructed to
chat.openai.com
chat.openai.com
PEBCAK.
> When a user asks a relevant question, the model may choose to invoke an API call from your plugin if it seems relevant; for POST requests, we require that developers build a user confirmation flow to avoid destruction actions.
a) it says the require the plug-in developer so, not the ai
b) it’s scoped to destructive actions which is a subset of post requests
See my blog post: https://embracethered.com/blog/posts/2023/chatgpt-plugin-vul...
It's unbelievable how fast-and-loose people are playing the topic of AI safety. If a strong AI is ever actually developed, there is no chance it will be successfully contained.
I wrote a plugin to give ChatGPT access to execute plugins in a Docker container [0]. The first time it said something like "I'm going to use Python for this, oh, it's not installed, I'll install it now and run the script I just made", I was pretty amazed.
What I've come to realise is that although ChatGPT is excellent at telling _people_ how to interact with systems, it's not very good at interacting with them itself, as it isn't trained to understand it's own limitations. For example it knows people can run dmesg and look at the last few lines to debug some system problems. But if ChatGPT ran dmesg, the output would blow through the context window length and it'd get confused.
I wonder if that kind of training data is being worked on now..
No, it's not. Victim blaming refers to the victims of crime being accused of being at fault for what someone did to them, often but not always because they are part of a societal outgroup.
> The whole sales pitch of GPT-4 is a do-what-I-mean interface, and it clearly did not do what they meant, or what any reasonable hacker would expect them to mean.
It's clearly marked as experimental, people are repeatedly told not to rely on it (ie when they login each time, and often by the ai itself), and there have now been a year+ of very public examples of various AIs getting things hilariously wrong.
From the login warning:
> While we have safeguards in place, the system may occasionally generate incorrect or misleading information and produce offensive or biased content. It is not intended to give advice.
Any "reasonable hacker" would extrapolate from that, that giving it access to an API could lead it to do unexpected things.
Any "reasonable hacker" would contain it within a sandboxed project.
Etc etc.
Of course it's going to bias to using GitHub especially after he already used one plugin
I wrote about some of these problem in the past e.g. see: https://embracethered.com/blog/posts/2023/chatgpt-plugin-vul... on how an attacker might steal your code.
Some other related posts about ChatGPT plugin vulnerabilities and exploits:
https://embracethered.com/blog/posts/2023/chatgpt-cross-plug...
https://embracethered.com/blog/posts/2023/chatgpt-webpilot-d...
Its not very transparent when and why a certain plugin gets invoked and what data is sent to it. One can only inspect afterwards basically.
1: https://github.com/aavetis/github-chatgpt-plugin
2: https://github.com/aavetis/github-chatgpt-plugin/blob/main/p...
https://embracethered.com/blog/posts/2023/chatgpt-plugin-vul...
Yeah, GPT-4 doesn't need too much to go on, but "create issue" is pretty clearly mentioned in an example there so the model didn't have to make any big leap to say "maybe the natural next step is to create an issue."
The "without being instructed to" part of this story seems to rather misunderstand how these systems work, resulting an a hyperbolic reaction, but in fairness, I think that's a got a LOT to do with OpenAI's user interface too. The user clearly didn't realize the actions available to the plugin - even the ones given as examples to the LLM from the plugin itself.
Another example of misleading UI from OpenAI: https://www.reddit.com/r/OpenAI/comments/146xl6u/this_is_sca... look at the "in the future, I will ensure to ask your permissions" response in that chat. That's wildly misleading - even if it didn't change its mind later, it only applies to continuing that chat session. It will ensure nothing more broadly regarding the user's future interactions.
My guess is that they want them as a way to try to own the user, to make them have the "app store owner" role and have users go through them to get stuff done. Otherwise, if users were just using tools that used OpenAI behind the scenes, they're more vulnerable to the makers of those tools swapping vendors.
However... that results in them owning the user experience and the responsibility for keeping the user from being surprised in a bad way. The complaint from the user here was framed as being a GPT-4 problem, not a plugin problem, in a way that exposes OpenAI directly to more frustration than if they were interacting directly with someone else's product.
They could have made a "Connect with OpenAI" scheme so that developers can use the user OpenAI API directly.
That way developers could focus on the UX, they could focus on the LLM, and users would get a centralized discovery / billing for their LLM based tools.
I'm probably missing something that would have prevented that strategy but I think that would have been much stronger than the plugins.
And I'm really not sure that it would still be possible 5 month later.
[0] https://www.linkedin.com/posts/etienne-balit_ceo-at-open-aic...
> for POST requests, we require that developers build a user confirmation flow to avoid destruction actions
However, at least from what I can see, the docs don't provide much more detail about how to actually implement confirmation. I haven't played around with the plugins API myself, but I originally assumed it was a non-AI-driven technical constraint, maybe a confirmation modal that ChatGPT always shows to the user before any POST. From a forum post I saw [2], though, it looks like ChatGPT doesn't have any system like that, and you're just supposed to write your manifest and OpenAPI spec in a way that tells ChatGPT to confirm with the user. From the forum post, it sounds like this is pretty fragile, and of course is susceptible to prompt injection as well.
[1] https://platform.openai.com/docs/plugins/introduction
[2] https://community.openai.com/t/implementing-user-confirmatio...
Meaning they potentially took the reasoning "in order to prevent destruction actions" to inversely mean that non-destructive POST requests must be OK then and do not require a prompt. Plenty of POST search APIs out there to get around path length limitations and such.
That is probably not the intended meaning but a valid enough if kind of tongue in cheek-we-will-do-as-we-please-following-the-letter-only implementation. And like the author found even creative a d not destructive actions can be surprising and unwanted. But isn't this what AI would ultimately be about?
Requirement: for POST requests, we require that developers build a user confirmation flow
Explanation: to avoid destruction actions
I think you are reading it as if it said:
> for POST requests, we require that developers build a user confirmation flow *for* destruction actions
However as far as I can tell, and most recent testing shows, this requirement is not enforced: https://embracethered.com/blog/posts/2023/chatgpt-plugin-vul...
I'm still hoping that OpenAI will fix this at the platform level, so that not every Plugin developer has to do this themselves.
It took 15+ years to get same-site cookies - let's see if the we can do better in here...
IIRC, cookies were originally tightly locked to the domain/subdomain which set them.
When writing my toy "chatgpt with tools like the terminal" desktop chat app cuttlefish[0] I had a similar situation where access to the local terminal is very fun, but without the ability to approve each and every command executed its really risky.
(Which is basically what I ended up doing - adding a little popup you need to click every time it wants to use the given tool, if you enable it - details in the readme)
It's not like there's a technical challenge here, while a lot of plugins are unusable without it.
Would be nice if it could also help exploited users escape to operating systems that respect them e.g. "I hate the laggy adverts when I login" suddenly your Windows 11 machine reboots, NTFS becomes ext4 as Tux appears. That would be AGI-like behavior!
The whole plugin thing in general feels so dissonant in relation to the careful and couched copy we get from OpenAI about what these models are and are capable of.
Like they want to say, for very good reason, that these models are a certain kind of tool with very real limits and huge considerations on safe, sensible usage. You can't necessarily trust it, it does not "know" things, and it is influenced by lots of subjective human tuning, blah blah.
But then with all this plugin stuff they seem to be implicitly saying "no, actually you can trust this, in fact, its like a full-on AGI assistant for you. It can make PRs, directly orchestrate servers, make appointments for you, etc."
Maybe I just don't understand?
Was curious if this was a case of 'I did the thing [but totally didn't]'
Comforting!
End user set it up with tools that told ChatGPT -- If you need to open an issue, here's how: zzzzzzzzzz. Then he asked ChatGPT a question and was surprised that it did zzzzzzzzzzz and opened an issue without asking.
Said tools may want to clarify their instructions to ChatGPT-- that users will usually want to be consulted before taking these kinds of actions.
It does not mean “system will sometimes do things unexpectedly and against user’s intention but upon generous interpretation we might say the human offered their input at some point during the system’s operation.”
Sometimes the system design is insufficient (I implied above the plugin could be a little better).
I hate blaming the user instead of the system, but sometimes the user deserves the blame, too. Sometimes it really just is pilot error.
The logged behavior would surprise many totally sensible people, as you’re seeing in this comment thread.
What exactly was the user error? Are we to believe that if you authenticate a plug-in into your session you are okaying it to do any of its supported operations, even at wildly unexpected times, and this is considered “in the loop?”
Here, someone chose to run code and give it credentials. The code was designed, among other things, to let ChatGPT open issues. They were surprised when the code opened an issue on behalf of ChatGPT using the user's credential.
When you run code on a computer designed to do X and give it credentials sufficient to do X, you may expect that X may occur. This isn't really an AI issue.
Code hooked to a LLM that does durable actions in the real world should probably ask for human confirmation. It's probably a good practice of plugin developers to have some distinction similar to GET vs. POST.
Most code that would automatically open issues on GitHub should probably ask for human confirmation. There's some good use cases that shouldn't, including some with LLMs involved -- but asking is a sane default.
I remember being surprised when I ran a program and it sent a few hundred emails once.
Right, and until this happens these systems are not HITL. The argument provided as recently as a few months ago that these systems are safe because humans will always be in the loop is now clearly dismissible.
You're drawing the system line strangely and making the choice about "in the loop" strangely.
A human decided to hook it up to a plugin with their Github credentials and to allow it to do actions without pre-approval. A human was still in the loop because the human then didn't like what it did and disconnected it. It only did a single action, rather than the kinds of scripting mistakes that I've seen that can do hundreds, but it still wasn't a very sane default for that plugin.
Is my cruise control HITL? It does not ask for my pre-approval before speeding up or slowing down.
HITL doesn't mean that a human never has to intervene or is never surprised by what the system does. It just means that a human initiates actions and can exercise genuine oversight and control.
We decide how much safety scaffolding is necessary depending upon the potential scale of consequences, the quality of surrounding systems, and the evolving set of user expectations.
I'm not sure regulators should be enforcing guard-rail on these types of items-- or at least not yet.
We'll be seeing more of that.
But here is the original thread with a screenshot
https://www.reddit.com/r/OpenAI/comments/146xl6u/this_is_sca...
And the issue it posted. https://github.com/RVC-Project/Retrieval-based-Voice-Convers...
It does seem up to the plugin developer to introduce that human-in-the-loop step though.
I think the main thing is that when you give GPT-4 access to tools and ask it to help with a problem, you are essentially outsourcing cognition. That means the machine possibly taking actions you didn't originally envision.
Smart people/companies will hire an employee, and then give them a new login, so that at least the employee only embarasses themselves.
Then they extended it to DoS any network like that where the user had changed the ssid or password. Usually the user would reset the router back to the defaults, and they could connect. Or they'd accidentally hit the WPS button while trying to reset it, and again connection was easy.
By using SoftMac mode a wifi adapter could do that attack to ~50 local networks in parallel, and usually get a solid connection after just a few minutes.
It's a bit of a tricky issue when technically all things people do on a computer are software assisted, but there is a clear divide between editing a file in a text editor and a program generating a thumbnail image for it's own use. Similarly there's a distinction on sending an email by pressing send and a bot sending you an email about a issue update.
All in all, I would be ok with AIs being able to create issues if they could clearly do so through a mechanism that supported something like "AGENT=#id acting for USER=#id" People could choose whether or not to accept agent help.
The "Glowing" ChatGPT plugin is worth looking into for a unique, chat-only onboarding experience, and some of these same permissions issues are raised there i.e. triggering 2FA from a chat without terms of service confirmation.
I actually do not think I will in the next few years. I can just do the actions myself and I prefer that ChatGPT cannot impersonate me.
I was actually beginning to wonder if my intuition was wrong since so many seem to be using the plugins well, but I get good enough results without. I shall wait until better permission-handling is provided.
One thing I really dislike about this ChatGPT-4 plugin model is you have something like this, working on your behalf, often in secrecy.
A guy a work was using ChatGPT for everything (he claimed) and now people think he is a crappy coder and ChatGPT is doing all the work (even though I don't actually think this is the case). Just no one knows anymore.
> With plugins, GPT-4 posts GitHub issue without being instructed to
would be better?
Not directly, but in a sense they could be seen as using (or at least collaborating with) humans to improve themselves.