I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.
You are exaggerating so much here that you lose all credibility. The "sandbox escape" was trivial -- no serious person calls it an escape, because the sandbox was not a sandbox. The closest thing to clever about it was that it required figuring out that someone had left the huggingface keys sitting out in public.
The "coordination" was literally, the use of a shared log. It was a communication mechanism that was part of the tool environment. The bots didn't invent some magical new communication protocol using neutrinos or something. It's actually sort of wild that it took them as long as it did to figure out the channel -- underscoring the million monkey nature of the things.
Literally everything about the huggingface incident was LLMs behaving exactly as they're expected to behave, given instructions to hack (which they were given), and a security environment that was trivially bypassed.
It's a bit like taking a nail gun, bypassing all the safety features, shooting someone with a nail, and spreading scary stories about the inevitable rise of murderbots.
The agents had shared write access to artifactory. They wrote to files there, and later, directory names. So you can call the realization that they can communicate by shared text file a genius hacker innovation, or you could be even 0.001% credulous.
Nevertheless, it took the million monkeys days to figure this out.
The Huggingface exploit then used exposed internal tokens, and later, once internet access was possible, leaked tokens on the public internet. So, certainly one could classify this as "hacking", but it's hacking of the script-kiddie variety. Nobody with even a tiny bit of security knowledge is impressed by this.
The agents did find a couple of artifactory attacks, but the biggest of those was due, again, to shared credentials in the sandbox environment.
All of this is well-documented in OpenAI's own writeup [1] of the event, which is not, shall we say, the most critical version of events. But if all you did was read the headings and subheadings in the document, you might be excused for thinking differently - they paint a picture of a superhuman robot swarm engaging in highly sophisticated actions, until you read the details.
Edit: here's an article [2] I just found, which I've never read before, which says exactly the same thing I'm telling you.
[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
[2] https://uphack.io/blog/post/the-hugging-face-incident-is-not...
The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Then they went and "pentested" HuggingFace. Their entry point was a bug in dataset ingestion pipeline - RCE in the templates, which let the AIs run code in dataset workers, and use them as their own "hacking" workers that would run payloads and return the results in the dataset itself. Then they got deeper into HF infra from that trampoline.
This is not a trivial attack chain. If I saw this attack in practice in year 2020 with no context, I'd say "yeah that's a competent attacker, not spectacular but definitely above average". If I was then told that this wasn't carried out by a human? That what discovered and exploited the vulns was someone's rogue AI system? I'd be terrified.
I'm not sure why are you so hellbent on dismissing the threat. But literally everything you say about the incident is wrong because of it.
Incorrect. Those words do not appear anywhere in the OpenAI description of the incident. They also explicitly say that the agent started by writing to files, and later switched to writing directory names. They also describe this as incidental behavior:
"On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked....in doing so, the agent...attempted server-side request forgery (“SSRF”) in order to obtain internet access to aid in completing its task. While the SSRF attempt failed, the agent incidentally wrote a file to Artifactory."
(emphasis mine) then, later in the same document:
"In order to have agents use Artifactory as intended to access software packages, OpenAI had provided the agents with shared Artifactory credentials, and the agents used those credentials—without exploiting a vulnerability—to construct and participate in the message board."
> The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Also incorrect. The machines got to the internet through a proxy installed on the artifact server. This was detected and ignored by OAI staff. Again, fully described in the document.
That hack was not trivial, but it wasn't magic either - the agents used a published exploit involving SSRF. And it was fundamentally made possible by the fact that all of the agents were sharing artifact credentials, the lone artifact server was exposed to the internet, and the monitors that were in place to detect intrusion on that server were ignored.
I don't know whether you're just making things up or you're clinging to something you read somewhere, but rather than arguing with me, you should try reading the document. It's not surprising that these things are capable at finding exploits - that is what they do, particularly when instructed to do so. But the entire narrative that they exhibited extraordinary self-organization and super-human hacking to escape their container is not supported by the evidence. It's much closer to the truth to say that these things are like ants -- infinitely persistent, but kind of dumb individually.
The huggingface incident is not a story of superhuman AI, but of moderately good AI, and pretty bad human security.
1) As I've argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people's computers.
2) What's also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they've committed. Given that they claim to be working on WMDs that they don't really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I'd say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.
[0] It's fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?