What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.
The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.
That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.
That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.
Orly?
Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.
Let's say the sandbox holds. It's a perfect, ideal sandbox! It's not even in the same universe as the rest of the internet. There's absolutely no way for the AI to escape!
Thus, "the unknown unreleased AI involved in the HuggingFace incident" doesn't actually hack HuggingFace. Because it can't! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes "GPT-6 Astra".
Then a web developer in Brazil gives his $100/mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that "GPT-6 Astra" is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it's a random developer in Brazil who gets blamed, and billed, and probably sued too.
You can't and shouldn't rely on a sandbox. An AI that's only safe if you keep it in the world's most ideal perfect sandbox is a disaster waiting to happen.
This might have gone okay if they weren't testing to see how well the tools attack computers, but, well, that's what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.
If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company's infrastructure, and then against another company is "we disabled the cyber classifer" and "we gave it an exploitation ability eval"?
AI is a dangerous technology.
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.
You are exaggerating so much here that you lose all credibility. The "sandbox escape" was trivial -- no serious person calls it an escape, because the sandbox was not a sandbox. The closest thing to clever about it was that it required figuring out that someone had left the huggingface keys sitting out in public.
The "coordination" was literally, the use of a shared log. It was a communication mechanism that was part of the tool environment. The bots didn't invent some magical new communication protocol using neutrinos or something. It's actually sort of wild that it took them as long as it did to figure out the channel -- underscoring the million monkey nature of the things.
Literally everything about the huggingface incident was LLMs behaving exactly as they're expected to behave, given instructions to hack (which they were given), and a security environment that was trivially bypassed.
It's a bit like taking a nail gun, bypassing all the safety features, shooting someone with a nail, and spreading scary stories about the inevitable rise of murderbots.
The agents had shared write access to artifactory. They wrote to files there, and later, directory names. So you can call the realization that they can communicate by shared text file a genius hacker innovation, or you could be even 0.001% credulous.
Nevertheless, it took the million monkeys days to figure this out.
The Huggingface exploit then used exposed internal tokens, and later, once internet access was possible, leaked tokens on the public internet. So, certainly one could classify this as "hacking", but it's hacking of the script-kiddie variety. Nobody with even a tiny bit of security knowledge is impressed by this.
The agents did find a couple of artifactory attacks, but the biggest of those was due, again, to shared credentials in the sandbox environment.
All of this is well-documented in OpenAI's own writeup [1] of the event, which is not, shall we say, the most critical version of events. But if all you did was read the headings and subheadings in the document, you might be excused for thinking differently - they paint a picture of a superhuman robot swarm engaging in highly sophisticated actions, until you read the details.
Edit: here's an article [2] I just found, which I've never read before, which says exactly the same thing I'm telling you.
[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
[2] https://uphack.io/blog/post/the-hugging-face-incident-is-not...
The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Then they went and "pentested" HuggingFace. Their entry point was a bug in dataset ingestion pipeline - RCE in the templates, which let the AIs run code in dataset workers, and use them as their own "hacking" workers that would run payloads and return the results in the dataset itself. Then they got deeper into HF infra from that trampoline.
This is not a trivial attack chain. If I saw this attack in practice in year 2020 with no context, I'd say "yeah that's a competent attacker, not spectacular but definitely above average". If I was then told that this wasn't carried out by a human? That what discovered and exploited the vulns was someone's rogue AI system? I'd be terrified.
I'm not sure why are you so hellbent on dismissing the threat. But literally everything you say about the incident is wrong because of it.
Incorrect. Those words do not appear anywhere in the OpenAI description of the incident. They also explicitly say that the agent started by writing to files, and later switched to writing directory names. They also describe this as incidental behavior:
"On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked....in doing so, the agent...attempted server-side request forgery (“SSRF”) in order to obtain internet access to aid in completing its task. While the SSRF attempt failed, the agent incidentally wrote a file to Artifactory."
(emphasis mine) then, later in the same document:
"In order to have agents use Artifactory as intended to access software packages, OpenAI had provided the agents with shared Artifactory credentials, and the agents used those credentials—without exploiting a vulnerability—to construct and participate in the message board."
> The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Also incorrect. The machines got to the internet through a proxy installed on the artifact server. This was detected and ignored by OAI staff. Again, fully described in the document.
That hack was not trivial, but it wasn't magic either - the agents used a published exploit involving SSRF. And it was fundamentally made possible by the fact that all of the agents were sharing artifact credentials, the lone artifact server was exposed to the internet, and the monitors that were in place to detect intrusion on that server were ignored.
I don't know whether you're just making things up or you're clinging to something you read somewhere, but rather than arguing with me, you should try reading the document. It's not surprising that these things are capable at finding exploits - that is what they do, particularly when instructed to do so. But the entire narrative that they exhibited extraordinary self-organization and super-human hacking to escape their container is not supported by the evidence. It's much closer to the truth to say that these things are like ants -- infinitely persistent, but kind of dumb individually.
The huggingface incident is not a story of superhuman AI, but of moderately good AI, and pretty bad human security.
1) As I've argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people's computers.
2) What's also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they've committed. Given that they claim to be working on WMDs that they don't really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I'd say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.
[0] It's fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?
They'll never let that happen because it would destroy their credibility.
It's like they're telling the world about this dangerous, possibly world-ending pathogen that they're developing, but they're evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.
We can (and at the time, did) mock Lemoine for his reasoning; nevertheless, he was willing to violate an NDA on the basis of what we would now call AI psychosis.
We are sill getting people resigning to talk about the risks as they see them. It would be great if, four years from now, we look back on them the same way we look back on Lemoine; but we can reasonably forecast that more staff will continue to say these things, especially as the public statements from several of the CEOs is compatible with "this has the potential to go very wrong, we should do something about that".
You may think the CEOs are lying*; but I think leaks from employees are what matters here, and regardless of if those statements are lies or not, they are statements which may increase such leaks.
* I mean, Musk did say "With AI we are summoning the demon"; given he now talks of a robot army and factories on the moon pumping out orders of magnitude more compute than humanity's current aggregate power use, either he's a liar or he's Faust. Or both.
The most recent DNS sandbox escape is similarly ridiculous.
E.g. even the basic thing (block the internet), was not actually properly blocked.
Some comments: https://techcrunch.com/2026/07/22/how-an-openais-human-mista...
That seems pretty locked-down to me. I don't think it's reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.
https://securityaffairs.com/195774/ai/openai-ai-models-explo...
It is reasonable to expect for companies to select third-party components that are fit for the purpose. Artifactory was not running in the sandbox, but rather as the edge, so it is in the sandbox'es trust boundary. Same sandboxing requirements would apply for this software too as it is pure dependency.
Security trust boundary was extended to include Artifactory as a dependency, but Artifactory was not fit for the job, and sandboxing failed. And as the network isolation was not good enough, the impact was catastrophic.
Seems pretty wild these things were hacking government websites etc but yeah no one picked that up until the victims reported it?
"“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” "
So sounds to me like they're saying: "We took off the guard rails to see how bad it could act and it acted bad".
Practical AI deployments aren't going to do ridiculous bullshit like "airgap the server farm" or "route all inputs through a data diode". They'll give an AI root access on production servers so that it can run diagnostics live during an incident. Then they'll forget to revoke that access.
If an AI can't be trusted not to take malicious actions in pursuit of its given goals even if deployed in the most half-assed manner and given more than enough access to take those malicious actions, we have a problem. Evidently, we have a problem.
Given the content of his recent interview with Ezra Klein, I wonder if we're going to see Hugging Face press charges against OpenAI. When the acquisition was publicly announced, I thought that a significant reason for the acquisition was to hush them up, but now I'm not so sure.
AI safety solved: driving humankind to extinction is now illegal! Skynet is outlawed! Rejoice!
How were the sandboxes poor?
It's not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.
Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.
No, we didn't know that and this is how you find out they're very good at hacking
HN's memory is so fickle. Just a few months ago almost no one here believed Mythos could actually be as good at hacking as the company claimed. This was a novel concept when the companies experienced these breakouts.
We knew models much weaker than mythos were good at hacking the problem they had was that when finding exploits they had too many false positives.
Either way, putting artifactory on the sandbox security boundary is obscene negligence. There is no reason to believe artifactory is secure.
But I guess OpenAI's security researchers acting surprised is part of some grand conspiracy to manage PR?
https://www.youtube.com/watch?v=1sd26pWhfmg
We already knew LLMs were capable of finding exploits like this.
The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.
What exactly? They seem to be trivial SSR/path traversal and input validation issues. Including misconfiguration. Nothing novel.
The question anyone versed in security would ask was why anyone thought artifactory was an acceptable security boundary. I would never assume artifactory was secure. It's like someone telling me there is a 0 day in a wordpress extension. So what?
Also I don't think Qemu is secure either because it's millions of lines of C and C++.
Firecracker I can trust to be secure because it's 70k lines of human audited Rust. I know there are multiple people that have a complete understanding of the firecracker codebase.
There was no real isolation because a part of the system that doesn't provide any isolation guarantees was bridged to the internet. The next version of Artifactory, which you definitely wouldn't audit before you rolled it out, could simply add a public API that sends requests out to the internet.
Such an innocent upstream change would be equally catastrophic for your security model, which should demonstrate why it's negligent to rely on undefined behavior to enforce your security policies.
This should frankly be obvious any operator entrusted with running dangerous and possibly malicious code. Even if you don't know what you're doing, any LLM would tell you that this is a really bad idea if you simply asked. Don't rely on spacebar heating [1] to keep humanity alive.