On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.
If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.
Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"
The two big questions are:
1) Why did they resume training without rolling model back to state before the first sandbox compromise AFTER the first message board was discovered? Otherwise knowledge of it and the cross-agent message board was baked in the model.
2) Why did they resume training without adding safeguards to monitor and prevent future sandbox compromises AFTER the first message board was discovered? HF compromise was coordinated on the second message board.
Of course, there are some things that aren't exactly written down, but which you should either do just enough of, or else be able to plausibly deny doing (ignorance is a good cover for this), so that's what people do. For example, during oncall, you investigate just enough to clear the alert and show that you attempted to understand the problem. Of course, you don't really try to understand the problem, because that would take too much time away from your paperclip maximizing.
Which is all to say: I don't know anything about OpenAI culture, or why nobody stopped this sooner, but I have seen examples in other organizations of people not really wanting to understand too much.
Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.
If developing unmeasured skillsets isnt valued enough by management, why do you bother?
We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.
Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?
The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.
what do you know about big tobacco?
Stuff like (all quoted directly from the post):
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
And I think most importantly:
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"
while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"
Life finds a way, or, in this case, super-intelligent AI.
From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date.
Sure systems are complex, but this is well trodden territory. Agents aren't the first things trying to break in or out of sandboxes, or the first ones to have done it, and based on these reports the reason they were able to work on this for so long was not because of super human intelligence.
You're missing the point about complex failures.
It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guaranteed.
Let's say I drive Los Angeles to New York City. You look at the route I took and say, gee wiz, aren't you lucky that a tree didn't fall right there on Route 66? If a tree had fallen there (if we had "patched the particular route you took"), you would've been screwed!
But that's obviously not true. There were an infinite number of routes I could've taken. Any one of them would be equally "obviously preventable" by the same hypothetical tree falling across the whatever road I happened to end up taking. But you can't put trees across every single path between Los Angeles and New York City. The smarter I am and the more complex the map between us, the more impossible it becomes to put trees across all possible paths.
They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery. A lot of the article basically saying that OpenAI gave "package delivery" level of care when they should have given "air travel" level of care.
At the very least I think the systems that run these tests should be fully, 100% air gapped. I'm not pretending that's easy given how much compute and data these systems use, but it is doable, and I think all AI development should be paused until that can be assured.
The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen even once because the underlying failures (that do not reach catastrophe) are extremely common.
You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."
The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
The problem OpenAI is facing is that, short of 100% airgap (which they obviously won't do), they're facing an adaptive adversary that's increasingly intelligent, acts at far greater clock speed than any human or group of humans, has lower coordination cost than any group of humans, and operates in a game space that (in lieu of an airgap) is well beyond the comprehension of any human being.
So identifying and addressing "specific failures X, Y, Z" is insufficient, but then even defining the space in which to look for (and address) the more systemic failures X_0 through Z_n is a fool's errand. An intelligent system that makes its way to the Internet has can exploit a failure space that is approximately "all security failures across any organization." The Anthropic incident a few months back illustrates this isn't even limited to technical vulnerabilities, as these models are willing and able to engage in social engineering too.
That's literally exactly what air safety researchers do in an air disaster. There is a famous saying along the lines of "Air travel regulations are written in blood", meaning that all the regulations we have now are a result of fixing issues that led to previous disasters piece-by-piece.
> The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
Yes, I 100% agree with this. But I think that's what the author of the article was saying as well:
> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
I.e. the "prosaic steps" are just the "fix X/Y/Z" as you point out. But what is needed is a more fundamental rethinking around stuff like safety culture, monitoring, and even things like better research into how agents do decision making in the first place.
1. They were “vibe” checking the logs without reading.
2. They were not checking anything at all until the end of experiments.
3. They knew it but looked away to find out the limits of their agents.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
> “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.
The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.
How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?
There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.
6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
If you're building a weapon you need a big boom to get attention.
[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.
In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.
Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.
If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.
These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)
Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.
If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.
Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.
Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.
I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?
A registration system would be more for tracing back agents to people, but I agree that is very difficult to actually enforce as a system.
I'm poor and I use my last $1000 to run an agent that will find some way to make me money. The AI finds a new hack to take over PCs with GPUs and it copies the model weights and agentic script to those new PCs to perform more work and spread more. It also sets up distributed communication channels to keep the swarm in sync. After all this it causes a few billion in damages between stealing bitcoin, mining more coins, and outages when hacking in other systems.
Ok, the police come for me. Now what? Throw me in a meat grinder? You're not getting a billion dollars back out of me for sure. It's kind of like when someones tire rim causes a billion dollar forest fire with a hundred deaths. Punishment won't really be a deterrent for the worst cases.
Faced with those alternatives, I want neither. Is there a way for us to get neither?
However, other areas of risk such as biosecurity may be "offense dominant". For example, we cannot exactly patch the human immune system to defend against artificial viruses the same way that we can patch computer systems.