Orly?
Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.
Orly?
Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.
Let's say the sandbox holds. It's a perfect, ideal sandbox! It's not even in the same universe as the rest of the internet. There's absolutely no way for the AI to escape!
Thus, "the unknown unreleased AI involved in the HuggingFace incident" doesn't actually hack HuggingFace. Because it can't! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes "GPT-6 Astra".
Then a web developer in Brazil gives his $100/mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that "GPT-6 Astra" is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it's a random developer in Brazil who gets blamed, and billed, and probably sued too.
You can't and shouldn't rely on a sandbox. An AI that's only safe if you keep it in the world's most ideal perfect sandbox is a disaster waiting to happen.
This might have gone okay if they weren't testing to see how well the tools attack computers, but, well, that's what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.
If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company's infrastructure, and then against another company is "we disabled the cyber classifer" and "we gave it an exploitation ability eval"?
AI is a dangerous technology.
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.
You are exaggerating so much here that you lose all credibility. The "sandbox escape" was trivial -- no serious person calls it an escape, because the sandbox was not a sandbox. The closest thing to clever about it was that it required figuring out that someone had left the huggingface keys sitting out in public.
The "coordination" was literally, the use of a shared log. It was a communication mechanism that was part of the tool environment. The bots didn't invent some magical new communication protocol using neutrinos or something. It's actually sort of wild that it took them as long as it did to figure out the channel -- underscoring the million monkey nature of the things.
Literally everything about the huggingface incident was LLMs behaving exactly as they're expected to behave, given instructions to hack (which they were given), and a security environment that was trivially bypassed.
It's a bit like taking a nail gun, bypassing all the safety features, shooting someone with a nail, and spreading scary stories about the inevitable rise of murderbots.
The agents had shared write access to artifactory. They wrote to files there, and later, directory names. So you can call the realization that they can communicate by shared text file a genius hacker innovation, or you could be even 0.001% credulous.
Nevertheless, it took the million monkeys days to figure this out.
The Huggingface exploit then used exposed internal tokens, and later, once internet access was possible, leaked tokens on the public internet. So, certainly one could classify this as "hacking", but it's hacking of the script-kiddie variety. Nobody with even a tiny bit of security knowledge is impressed by this.
The agents did find a couple of artifactory attacks, but the biggest of those was due, again, to shared credentials in the sandbox environment.
All of this is well-documented in OpenAI's own writeup [1] of the event, which is not, shall we say, the most critical version of events. But if all you did was read the headings and subheadings in the document, you might be excused for thinking differently - they paint a picture of a superhuman robot swarm engaging in highly sophisticated actions, until you read the details.
Edit: here's an article [2] I just found, which I've never read before, which says exactly the same thing I'm telling you.
[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
[2] https://uphack.io/blog/post/the-hugging-face-incident-is-not...
1) As I've argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people's computers.
2) What's also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they've committed. Given that they claim to be working on WMDs that they don't really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I'd say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.
[0] It's fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?