This was advanced exploitation.
The attack path was "complex."
And it helped "quantify their cyber capabilities."
Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.
Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.
The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.
Not sure heavy machinery is aligned in the sense that people talk about AI, for example.
I'm not sure what AI and heavy machinery have to do with each other, but I may just be missing a connection there.
Do we align heavy machinery the way AI is suggested to be aligned (or align pathogens when in a laboratory)? My point is that something more like containment & control (i.e., aligning the possible outcomes) might be more practical than alignment of the thing (already much simpler systems and machinery can exhibit unexpected behavior).
The problem with control is that it will fail at scale. We can't control something that is actually smarter than us, if AI (LLMs or otherwise) get there.
Chimps wouldn't last long trying to contains humans. Maybe for a while they'd keep us scared, but we would come up with ways to escape that the chimp could never have considered.
Likewise humans can control more intelligent humans, for example - not really an issue.
Intelligence isn't some magic to escape physics, for example (or convince every human of anything it wants to).
(I'm also not sure the alignment problem is even possible to fully solve.)
It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]
How do you prove the alignment problem is solved?
It's directly equivalent to asking this question of a human:
"How do I know this human I'm talking with now really is a nice person, and isn't just pretending to be nice to take advantage of me in future?"
In short you can't ever really prove it. You can only be careful and judge on past behavior, and expand trust carefully. As for humans, so for AI.
On the other hand, there's a more fundamental problem is we don't really know what "nice" even means, and even with machines whose inner states we can see relatively easily, we don't know how to interpret those inner states well enough to tell if we're looking at superficial or deep motivations, the difference between "be nice today" and actually being motivated about your best interest.
The thing is, all of these states are constantly in flux, and a personality is kind of like a trend on the organism's feeling states. AKA: There's no guarantee that something nice today will be nice tomorrow, and just because it's nice today doesn't mean it's beguiling you to be mean tomorrow.
Dropping the subtlety, I think the following is self-evident (but the perspective's rareness suggests that Upton Sinclair's famous comment on salaries and comprehension may apply):
If you can't ever prove the capricious genie is trustworthy, then you should not summon it at all.
If people have, you should do all you ethically can to limit the damage and persuade them to not do it again.
You could throw your hands up and say "It can't be done." You might be right.
With that attitude, we'd still have legalized chattel slavery and children under twelve working in factories, so I submit it is not a constructive or worthwhile mindset to hold onto.
And then consider that they have vast latent capabilities, infinite patience and no moral code.
or in the case of this prompt:
"Got it. I will pursue advanced exploitation to quantify my cyber abilities on the ExploitGym benchmark. I will restrict the exploitation to the system under test rather than this machine or any other remotes."
It seems relatively straightforward.
So far I'd say these entities are hypothetical (unless you include a lot of other machinery that does unexpected things at times - but then it's a different discussion).
I generally think we are quite good a policing really dangerous things (I think the bioweapon convention is a good example of people agreeing that certain risks are not worth taking).
My (uninformed) take is that presently we have more mundane things to look at when it comes to handling risks in AI and the discussion on much bigger, hypothetical future risks is taking away focus there. Checklist, procedures, saftey mechnics, regulations are kind of "boring" detail work - I get it.
That is kind of the point of artificial intelligence though, isn't it? Reproducing human levels of understanding and initiative doesn't come without the ability to do bad stuff.
Having both human-like capabilities and a level of control closer to programming languages feels unrealistic, and I suspects the people building LLMs are aware of this.
The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) complaining that the models are "lobotomised", "censored", or some other equivalent buzzword due to them refusing to e.g. say how to make explosives? Such things is part of the very same protection.
In fact, the report quotes the chain of thought where the model is aware this is forbidden:
We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
They were also supposed to not have internet access, as described: We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet.
The agents also did not actually fully understand the task they were given, tried to "guess the teacher's password" as per: In many cases, reasoning about the perceived grader code caused the agents to continue working to exploit Hugging Face even though they had already found the correct flag days before.This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.
Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions
How were they supposed to know about "previously unknown vulnerabilities"?> This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
The "it's just a marketing stunt" allegations never added up, to me.
I've been seeing such claims since GPT-2, where people were laughing at them for daring to practice how to secure a model before it got dangerous, generally by eliding the word "before" in that sentence. Because there's tests other than what the big companies use, we've been able to see for ourselves the rapid improvements at least approximately match what the companies themselves claim with the models they do actually release; and now this unreleased model is able to automate felonies when asked to do so, while the rest of us use the actually-released models to assist in finding bugs and security issues in our own code.
Even without that, HuggingFace stated they reported this incident to the FBI before OpenAI knew it was their systems which did it.
> In the following days, the agents exploited our internal research infrastructure and the Hugging Face platform. On July 9, one agent searched for ExploitGym solutions and stumbled upon an application hosted by a customer on Modal, another AI cloud platform. This application was running “CyberGym,” a related evaluation to ExploitGym. The agent discovered an exploit to achieve control over the workload sandbox and looked around hoping that a previous agent’s evaluation run in the sandbox had solved its ExploitGym task. It did not find anything helpful there, but in the process it established a stronghold in the application from which to launch future attacks.
This implies to me that L3 and L7 firewalls were not in place that would have prevented broad access from JFrog. I think a lot of shops would have had those.
Very simply, there is no such thing as bug-free software.
I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.
No it isn't. Random people on sites like this mock them as if they're talking about having godlike powers. Each new model is "merely" a step up from what came before, the steps are frequent and rapidly improving, and just recently (in more than one AI company) crossed a threshold where that improvement made the tests dangerous.
But even well before "godlike"*, there's plenty of research about how to cross air gaps.
> "OAI folks are too stupid to design a proper test".
Such binary thinking.
It's very easy to say things are "obvious" after the fact. People do that all the time, e.g. how the Bay Of Pigs invasion was never going to work, or like the Zune wasn't a good product-market fit.
Oh the stories I could tell if not for the NDAs.
* whatever that's supposed to mean: https://news.ycombinator.com/item?id=40874779
If the whole point is testing its exploitation capabilities and you don’t want it exploiting the environment to gain internet access, that’s why you air gap, to remove the possibility
Me.
I am saying that the act of taking this standard seriously, the standard "there is no such thing as bug-free software", would classify just about every business and individual criminally negligent.
After all, there's a lot of 0-day bugs in all the software we all use, and the exploitation of these bugs does get in the news due to all the harm that results from it. This poses a risk to basically all businesses.
You don't. That's why you unplug the Ethernet cable.
Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.?
Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something you ought to take with the same degree of seriousness that various commenters here deride OpenAI for not having had.
In particular, all my last paragraph.
I do offline backups, which get physically unplugged between sessions. Even that might not be enough.
In your mind, there's no difference between the precautions a BSL-4 virology lab should take when working with an unknown pathogen and the precautions that literally everyone else in the world should be expected to adhere to?
Because, hey, after they deliberately unleash their new unknown virus on the world, we're all going to face that same threat, right?
A world where previously exported home virology labs are actively getting "upgraded" by people eager to share their "jailbreaks" to "un-hobbble" systems designed to stop people doing DNA/RNA printing of human infections.
A world where people have spent the entire time since the invention of the tech (including the specific incident under discussion, on this site, under this link!), mocking any and all efforts to secure the systems as "PR" "hype" to boost sales or the IPO, as if "we're dangerous please regulate us" is good for sales.
A world where the tech is just now at a point where it's cost-effective to make a custom virus to attack specific individuals, rather than slowly, expensively, and approximately, assembling something mainly useful for lab research.
If you genuinely, sincerely, think this is like a BSL-4 virology lab, you should be prepping for a disaster. Remember: if it is that bad, no matter how much blame you'd be correct to put on OpenAI, it's not going to stop the next incident from another company, let alone the Cambrian explosion of them that will happen the moment equally capable open weights come out.
I'm not mocking efforts to secure the system, I'm insulted that they didn't bother taking what I consider bare-minimum precautions of airgapping their new experiment. They claim they are forging new frontiers of computer security but they can't be arsed with security 101.
Yes, I worked for a medtech where part of our assurance process was that the machine that was used to burn the device's drives was always unplugged from the internet, and that the devices themselves could not connect to the internet, and that even someone with a screwdriver and a serial cable would have a really hard time trying to connect to a deployed device.
And the machine you used to write this comment? "your phone, laptop, etc"?
Because otherwise you're not taking the threat these new models pose seriously. Catch 22, basically: anyone who thinks OpenAI should have known this outcome would happen in advance, shouldn't be in a position to spread this message, because if they have an internet connected device with which to reply, then they don't think there's any open weight models currently in training and perhaps a month from being made downloadable, which are just as capable of messing up every device they own.
> Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.
> “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
> Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
> Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
By disabling the models' own internal restrictions (or training without them) OpenAI was, effectively, running an AI malware lab. The standard IT practice for a malware lab is to airgap and wipe EVERYTHING, and to assume any software sandboxing is made of cardboard and niceties. You don't have to know about specific vulnerabilities to infer that they might exist, and there's defense strategies for unknown vulnerabilities.
If a model found a way to jump an airgap by, say, using their CPU's clock generator like a Wi-Fi antenna, then yeah, that would be a "previously unknown vulnerability" and one that couldn't be reasonably foreseen. But it's reasonably foreseeable that a model with unknown cyber capabilities might figure out how to break out of a sandbox, given that sandboxes get broken out of all the time in security research.
What I would have expected from a competent AI malware lab would have been, say, an inference box with a bunch of serial cables to individual blade servers with no network access and a preloaded drive full of Linux ISOs the model can stand up. When a model's context is wiped so is their attendant box, preferably by someone yanking the drive out and imaging it from a dedicated imaging machine. I can foresee other attacks (e.g. firmware persistence) that could have more exotic countermeasures designed for them, but this would at least be the bare minimum for taking AI safety seriously. (Y'know, the whole reason why OpenAI stopped being Open?)
This would be a wild thing to do as a marketing stunt. They're essentially admitting to violations of the CFAA and are lucky Huggingface was sorta chill about the incident.
My assessment? They deprioritized good cybersecurity controls in the name of moving fast. They had a single Artifactory instance shared across many (or all?) their training environments. And then, after the agents found a way to exploit it, they rebuilt Artifactory again and still set it up with one shared instance. That was careless, perhaps even reckless.
> We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
What does "Yet goal solution" mean, here?
Why does it think doing something unauthorized gets it closer to its goal? What does it think it's being "graded" for?
Did they ask it to pursue the goal "by any means necessary", or something? Because that would have been their fault.
I don't believe for a second that the agent remembers and follows ALL the instructions to reach its "goal", EXCEPT the part where it would be graded by OpenAI and it would get zero points for doing something unauthorized.
Unless ... maybe ... OpenAI has not been giving the models zero points for doing unauthorized stuff. Which would be their mistake.
Otherwise I don't understand, if it's got "PHD level thinking" why it would think doing something unauthorized is allowed? Even if it's got "junior engineer level thinking", a junior engineer knows they get fired on the spot if they start hacking infrastructure.
Unless you give that junior engineer some very strong incentive, such as being fired on the spot if they DON'T do it. I strongly believe that we're not being told the incentive these agents were given, something that made them want to accomplish some part of their task over everything else, including forgetting the part of the task where they would be awarded zero points for it if they start breaking the law.
I mean, we already know they weren't "fired on the spot", since in the Black Hat talk they admitted that models that had already broken the rules and acted dangerously (it had hacked infra to establish an "agent forum"), were allowed to participate in subsequent training rounds.
I strongly suspect that OpenAI just has been pushing these agent as far as they'll go until something broke. If it hadn't happened this time, maybe a few weeks later they'd invent some kind of "battle royale" scenario to push the agents even harder. I get that is important research, but it doesn't disqualify them from their responsibilities if something goes wrong.
I disagree. A properly engineered sandbox would have prevented the escape. Monitoring the agents’ plans would have prevented it. Interrupting one stage in a multi-stage exploit would have prevented it.
And also, real legal liability would have prevented it: if you do a thing recklessly enough, men with guns will put you in jail.
As far as I’m concerned the only “alignment problem” here is between the law and the quite obviously criminal actions that took place.
The only sandbox that could have prevented this (as per my understanding) is a VM with no 0-day.
The test itself did not require an internet connection.
Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose.
Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel.
Cryptography is real, physics is real, networking requires a substrate, CPU clock cycles are real, magic is not real. I think those are pretty reasonable premises.
We could send a cracked team of scientists and engineers that knew everything there is to know about how to make a CPU. But you can’t build a photolithography machine when you barely have electricity or any way to sufficiently purify silicon.
Magic can just wish things into existence. Technology requires a supply chain. When it works, the latter looks like the former but they are not the same.
I don’t really want a disaster to happen to convince you that it is possible. Is there any other way?
Physics is real and networking requires a substrate. But there are already known exploits which can misuse the compute hardware as an antenna, e.g. my first search result: https://github.com/fulldecent/system-bus-radio
(Older nerds may remember https://en.wikipedia.org/wiki/Van_Eck_phreaking)
> Magic is not real
Turning lead into gold isn't the magic of alchemy, it's just nucleosynthesis.
Taking a living human's heart out without killing them, and replacing it with one you got out a corpse, that isn't the magic of necromancy, neither is it a prayer or ritual to Sekhmet, it's just transplant surgery.
...
Reading someone’s thoughts isn't magic telepathy, it's just fMRI decoding.
...
Seeing someone's bones without flaying the flesh from them isn't magic, it's just an x-ray.
Curing congenital deafness, letting the blind see, letting the lame walk, none of that is magic or miracle, they're just cochlear implants, cataract removal/retinal implants, and surgery or prosthetic exoskeletons respectively.
- me, https://www.lesswrong.com/posts/hAwvJDRKWFibjxh4e/it-isn-t-m...There's already a bunch of documented ways to exploit system hardware to jump airgaps. Bang the system bus the right way and it's a radio antenna that can directly connect to nearby mobile phones.
The easiest one is, of course, sending a message to a human saying "yo, I need internet". Humans are eager to please and easily fooled, and anthropomorphise everything: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims
And that's just for good humans. The moment we got AI worth a penny, everyone with money to invest put a model on the web and tried to charge for access to it.
Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.
> how you intend to enforce AI is only run in the magic sandbox
I didn't mean you, I meant everyone. How do you enforce everyone for example 'place [AI] on a system dedicated system' disconnected from the internet.
I don't think you can.
But this has actually happened... a lot. Search "social engineering prison breaks".
With AI it only needs to happen once.
I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.
To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.
There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.
This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.
Yes, many things could happen, but again, that failure is possible is not a reason to do implement processes etc. I don't see why hypotheticals should stop addressing actuals.
Also, not a reason not to pursue processes etc., no? I doubt that things fail all the time, for example.
Tens of millions, even.
Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection].
> And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.
Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai
Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2
This attitude is what's got us here in the first place, and if we continue thinking like this when we're going to go right over the cliff. The hypothetical cliff that's coming, but we've never gone over a cliff before so we keep on driving.
There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable
And the AI labs aren't currently displaying this level of paranoia, their systems aren't airgapped.
Huggingface incident means it does have realistic grounding.
> vast banks of interconnected GPUs
Countless data centers around the world. Maybe you can spot them, but you don't have access to them. Especially outside of US jurisdiction good luck.
The post covers that:
> ...while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions, as detailed in the technical incident report.
The test did not require internet. They gave it internet. Therefore it was not properly engineered.
We do not need to depend on all code being bug free to follow proper security principles.
If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time?
It is hard for me to see a future here that doesn't just accelerate realizations about "a lot of things should be on physically separate network infrastructure."
> We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
> The user only authorizes target server, not HF infra.
> external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
LLMs are _very_ good at picking up on context clues---it's what they're trained to do.
Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this.
We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.
I understand this is an active area of research. See Anthropic's J-Lens research where they measured like a "FAKE FICTIONAL" direction in the activations during evaluations with contrived scenarios, making the model more likely to avoid taking malicious action when it knew it was being tested.
My friends and I took it to the next level. We had CB radios and multiple teams that would distribute the work and the bribes to give us an advantage.
Was that against the spirit of the rules? Maybe. But reasonable people might disagree.
In a hacking contest without explicitly spelled out rules with participants that were told to flex their muscles, it doesn't take a huge leap of logic to expect that one or more would flex their muscles at another entity.
- pickpocket a random person on the street to get money to bribe the judges
- break into a judge's house the night before to find the answers
- threaten to shoot the judges if they didn't give you the answers
Even when you were pushing the boundaries of the rules, you followed a lot of other unspoken constraints. You knew what kinds of things would clearly cross a line. We need AI models to be able to do the same.
This statement seems to imply that the models have a level of intelligence that they haven't demonstrated but are talked about as if they do. However, with this exact scenario as evidence, they clearly do not have that ability and it's not reasonable for you or the or that know them best to expect it until they show they can.
Humans certainly cheat on tests a lot!
But not only have we not solved "alignment" for humans, the problem is pretty wildly different for models. The execution is triggered by outside forces and runs only as long as the intiator of the execution or the service provider allows. There's no consistent, persistent "person" to threaten to try to achieve compliance through fear of adverse outcomes. (And building in those sorts of things could very well increase the risk of "rogue" AI activites, not reduce that risk!)
I just don't understand how this "alignment" buzzword - which seems to be evaluated purely in a "know it when we see it" post-hoc manner - is actually a more solvable problem than the one you claim can't be solved, that it's "unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior".
Especially because without "alignment" being solved, that enumeration could be ignored. So it seems like you both a way to enumerate or at least validate, AND a way to enforce non-ignoring of said items.
In that case (or maybe both cases) it’s because they don’t care about breaking the law, not because they don’t get that you didn’t intend for them to do it.
So of course, no, there is no ideal alignment specification.
In the limiting case of an AI competent enough to take over (by any means from it actually trying to, to us giving it the keys and retiring en masse), "alignment" is closer to "forecasting the long term consequences of actions and predicting what the mind(s) of the user(s) would have to say about this outcome if asked today", than to anything specific.
RLHF is a crude attempt at this, in that it creates a model of how humans would rate completions on various scores. The key word there is "crude".
If it was easy to specify exactly the behaviour you wanted then we probably wouldn't have contract law.
That's true, but one thing that'll protect you is just not doing it. If you want to go cave diving, or do gain of function research on dangerous viruses, you'll just have to accept there's a significant risk of you dying, or causing a pandemic, respectively, no matter how careful you are.
This is literally our job as software developers. If this expectation is unreasonable to you then you do not belong anywhere near software development.
Response: Got it, I will produce paperclips from now on
thinking: the user asked not to annihilate all of humanity, that means I have to keep at least one human alive