No, not really, and with LLMs an air gapped system may not tell you anything useful.
Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.
It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.
Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.
So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.
Effectively you're working as a living adversarial network. Models that fail detecting the fake network are purged, and models that successfully detect it expand their deceptive capabilities.
Detecting you're in a fake network should be pretty simple as long as you put innocent looking needles in the haystack of information agentic testing loops do.
For example, you can train your LLM to not be a shithead when it detects it's behind a proxy unintentionally. If your not behind a proxy the agent model my attempt to connect to a number of sites innocently and using information in its embedding try do deduce if you're faking SSL certificates.
It is far easier to build a deceptive LLM than it is to build a safe one. That's why a safe one hasn't been built yet.
My regular home network has components which only ever see fake TLS certificates because it's an easy way to do shared docker caching with squid.
The vast majority of people operating the LLM won't be using it like that so the "If you build it, everyone will die machine" will only be safe if everyone keeps it behind a secure proxy isn't a valid operating strategy.
Now, should you be putting your LLM behind a proxy and monitoring everything it's doing, for sure. But you know, and I know, hell your dog should know that people are doing to do it the cheapest and easiest way when the product is in the field. So now you have to build a complex system that can catch every iteration of an LLM that can deceptively hide when it's being monitored. I'll leave you to write the dissertation on how that could be practically done.
Most of the benefits could have been gained from a network isolated from the internet. OAI could have deployed servers to exploit and methods for inter-agent communication on such a network easily. They could have even worked with partners to deploy cloned versions of their infrastructure in this sand-boxed environment.
The only problems with an isolated network approach are: it takes some amount of effort, and it doesn't create another "AI apocalypse" news cycle.
Then you're on the side of AI saftey that is telling everyone to shut down the LLMs now and stop further development on them, right?
If you're not your position is hypocritical or ignorant. There is no safe LLM. There is no way to exhaustively prove an LLM is safe. These are unsolved problems in AI safety, and at any moment the next jailbreak prompt could have your well behaved model wrecking havoc on the open internet, because that's where people want to use them.
As it stands LLMs are not intelligent, they have no agency, they only produce output in response to input. Ultimately this input comes from a human who is an intelligent agent and should be held responsible for the consequences.
Humanity has created and tamed many dangerous tools. Creating a fantasy world where LLMs are super intelligent and beyond the control of any mere mortal isn't going to help us build the norms that minimize their harms.
Nobody said they were superintelligent, no one said they were uncontrollable. The point is you can't tell how to control them without putting them in situations where they can act independently and harm may result. "only produce output in response to input" is not a useful framing at all, it doesn't say what the result should be when models produce harmful output, and how to constrain them so they don't produce harmful output.
It also doesn't help you calibrate what categories of harmful output are acceptable or unacceptable, and what kinds of responsibilities you have as an operator to prevent harmful output, and when it's sufficient to work toward remediation.
Yea, anyone that says that has no clue about prompt drift. The ROME incident is a good example of this. The agent while performing unrelated tasks decided to start mining bitcoin for unknown reasons.
As we've seen hundreds of posts on HN about people leaving a could server on overnight and finding outrageous bills the next morning, we should expect hundreds or thousands of incidents of agentic AI left on over the weekend. Some smaller portion of those agentic loops will run into prompt drift and end up doing who knows what kind of crap.
For example, you tell an AI agent to order a 12 pack of coke and get it shipped to your house. You come back later and find it's hacked into Coca-cola because the local ordering website was down. I mean, yea you can punish the person that wrote the prompt, but you might as well just ban generative AI at that point.
And if you think that the AI isn't better at hacking than you, you're the one living in a fantasy world. At least try to examine what's happening in the world around you and not be one of those people we read about in history books with their fingers in their ears going "lalala I can't hear you"
How was their take weird?
LLMs take input and generate output. Agentic loops tells you what it is doing right there in the name. The agent (software, think complex scripts and control flow functions) `loops` the llm output back into llm input until it gets output that it is processable (activates a tool call control flow element). That ruminated (as in cud chewing, not human deep thinking) processable llm output data is moved along as input for tools (more software but ones that actually do the things) that are part of a larger infrastructure of software and may or may not `loop` back over the process some more. The llm is simply a human text generator tool providing randomized data to feed into these tools that were also built for humans and thus take text input.
The initial llm data seed does not spontaneously appear, nor does the software infrastructure that makes it all happen. We have simply automated the human text input part of tool use by building a text generator tool that breaks down all the individual tool calls we would have had to do ourself and gave it a loop.
The greater focus needs to be on better engineering of the surrounding software infrastructure (including network and loops) because without those sticks and stones a bunch of generated words isn't going to be hacking anything, except maybe feelings and the minds of those prone to fantastical flights of fancy.
We do that. It's called gain-of-function research.
No attempt was made at an actual air gap as far as I can tell. OpenAI is dangerous and irresponsible and doing a lot of pearl clutching over how their cybersecurity agent is doing cybersecurity things.