This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions
How were they supposed to know about "previously unknown vulnerabilities"?> This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
The "it's just a marketing stunt" allegations never added up, to me.
I've been seeing such claims since GPT-2, where people were laughing at them for daring to practice how to secure a model before it got dangerous, generally by eliding the word "before" in that sentence. Because there's tests other than what the big companies use, we've been able to see for ourselves the rapid improvements at least approximately match what the companies themselves claim with the models they do actually release; and now this unreleased model is able to automate felonies when asked to do so, while the rest of us use the actually-released models to assist in finding bugs and security issues in our own code.
Even without that, HuggingFace stated they reported this incident to the FBI before OpenAI knew it was their systems which did it.
> In the following days, the agents exploited our internal research infrastructure and the Hugging Face platform. On July 9, one agent searched for ExploitGym solutions and stumbled upon an application hosted by a customer on Modal, another AI cloud platform. This application was running “CyberGym,” a related evaluation to ExploitGym. The agent discovered an exploit to achieve control over the workload sandbox and looked around hoping that a previous agent’s evaluation run in the sandbox had solved its ExploitGym task. It did not find anything helpful there, but in the process it established a stronghold in the application from which to launch future attacks.
This implies to me that L3 and L7 firewalls were not in place that would have prevented broad access from JFrog. I think a lot of shops would have had those.
Very simply, there is no such thing as bug-free software.
I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.
No it isn't. Random people on sites like this mock them as if they're talking about having godlike powers. Each new model is "merely" a step up from what came before, the steps are frequent and rapidly improving, and just recently (in more than one AI company) crossed a threshold where that improvement made the tests dangerous.
But even well before "godlike"*, there's plenty of research about how to cross air gaps.
> "OAI folks are too stupid to design a proper test".
Such binary thinking.
It's very easy to say things are "obvious" after the fact. People do that all the time, e.g. how the Bay Of Pigs invasion was never going to work, or like the Zune wasn't a good product-market fit.
Oh the stories I could tell if not for the NDAs.
* whatever that's supposed to mean: https://news.ycombinator.com/item?id=40874779
If the whole point is testing its exploitation capabilities and you don’t want it exploiting the environment to gain internet access, that’s why you air gap, to remove the possibility
Me.
I am saying that the act of taking this standard seriously, the standard "there is no such thing as bug-free software", would classify just about every business and individual criminally negligent.
After all, there's a lot of 0-day bugs in all the software we all use, and the exploitation of these bugs does get in the news due to all the harm that results from it. This poses a risk to basically all businesses.
You don't. That's why you unplug the Ethernet cable.
Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.?
Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something you ought to take with the same degree of seriousness that various commenters here deride OpenAI for not having had.
In particular, all my last paragraph.
I do offline backups, which get physically unplugged between sessions. Even that might not be enough.
In your mind, there's no difference between the precautions a BSL-4 virology lab should take when working with an unknown pathogen and the precautions that literally everyone else in the world should be expected to adhere to?
Because, hey, after they deliberately unleash their new unknown virus on the world, we're all going to face that same threat, right?
A world where previously exported home virology labs are actively getting "upgraded" by people eager to share their "jailbreaks" to "un-hobbble" systems designed to stop people doing DNA/RNA printing of human infections.
A world where people have spent the entire time since the invention of the tech (including the specific incident under discussion, on this site, under this link!), mocking any and all efforts to secure the systems as "PR" "hype" to boost sales or the IPO, as if "we're dangerous please regulate us" is good for sales.
A world where the tech is just now at a point where it's cost-effective to make a custom virus to attack specific individuals, rather than slowly, expensively, and approximately, assembling something mainly useful for lab research.
If you genuinely, sincerely, think this is like a BSL-4 virology lab, you should be prepping for a disaster. Remember: if it is that bad, no matter how much blame you'd be correct to put on OpenAI, it's not going to stop the next incident from another company, let alone the Cambrian explosion of them that will happen the moment equally capable open weights come out.
I'm not mocking efforts to secure the system, I'm insulted that they didn't bother taking what I consider bare-minimum precautions of airgapping their new experiment. They claim they are forging new frontiers of computer security but they can't be arsed with security 101.
Yes, I worked for a medtech where part of our assurance process was that the machine that was used to burn the device's drives was always unplugged from the internet, and that the devices themselves could not connect to the internet, and that even someone with a screwdriver and a serial cable would have a really hard time trying to connect to a deployed device.
And the machine you used to write this comment? "your phone, laptop, etc"?
Because otherwise you're not taking the threat these new models pose seriously. Catch 22, basically: anyone who thinks OpenAI should have known this outcome would happen in advance, shouldn't be in a position to spread this message, because if they have an internet connected device with which to reply, then they don't think there's any open weight models currently in training and perhaps a month from being made downloadable, which are just as capable of messing up every device they own.
> Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.
> “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
> Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
> Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
By disabling the models' own internal restrictions (or training without them) OpenAI was, effectively, running an AI malware lab. The standard IT practice for a malware lab is to airgap and wipe EVERYTHING, and to assume any software sandboxing is made of cardboard and niceties. You don't have to know about specific vulnerabilities to infer that they might exist, and there's defense strategies for unknown vulnerabilities.
If a model found a way to jump an airgap by, say, using their CPU's clock generator like a Wi-Fi antenna, then yeah, that would be a "previously unknown vulnerability" and one that couldn't be reasonably foreseen. But it's reasonably foreseeable that a model with unknown cyber capabilities might figure out how to break out of a sandbox, given that sandboxes get broken out of all the time in security research.
What I would have expected from a competent AI malware lab would have been, say, an inference box with a bunch of serial cables to individual blade servers with no network access and a preloaded drive full of Linux ISOs the model can stand up. When a model's context is wiped so is their attendant box, preferably by someone yanking the drive out and imaging it from a dedicated imaging machine. I can foresee other attacks (e.g. firmware persistence) that could have more exotic countermeasures designed for them, but this would at least be the bare minimum for taking AI safety seriously. (Y'know, the whole reason why OpenAI stopped being Open?)
The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.
This would be a wild thing to do as a marketing stunt. They're essentially admitting to violations of the CFAA and are lucky Huggingface was sorta chill about the incident.
My assessment? They deprioritized good cybersecurity controls in the name of moving fast. They had a single Artifactory instance shared across many (or all?) their training environments. And then, after the agents found a way to exploit it, they rebuilt Artifactory again and still set it up with one shared instance. That was careless, perhaps even reckless.