The agent harness is a process, like any other process in an OS.
You are a researcher running thousands of unattended automations that can hack a website without supervision. The first thing anybody will do is put security at various levels and isolate the network as much as possible. If something escapes your allow list, it should stop the processes as soon as possible.
You cannot foresee a bug in a server (like the Artifactory server in the Hugging Face incident). But you can isolate that server at the network level in the first place. So even if you give that server read-only access, no unexpected packets go out. It's not rocket science; it's something a billion-dollar company experimenting with what they promote as the biggest possible threat to humanity (if they do not handle it) could easily do.
They minimize their liability by changing the message to “oh look how powerful our models are, now we are going to have a public awareness report of the model deviations”. The message should be, “Sorry, we ran experiments without proper sandboxing; it’s our fault, and we changed our testing practices since then.” The former message puts all the blame on the smart, uncontrollable force of AI; the latter is what really happened: an irresponsible test over the Internet.
My argument is very similar to the article in the parent post:
The messages OpenAI published around the recent incidents emphasized their model capabilities but shifted away from their negligence in how they set up and monitor their evaluations.
Sort of like the “guns kill people” vs “people kill people” debate.
Deliberate wording to minimize perceived culpability for the agents actions.
Now do I think that’s the reason? It certainly isn’t a new thing for companies to try to do that. Shift blame that is.
Regarding civil liability, I’m not making that argument here. But it makes sense from a public perception viewpoint why they would want the agents to appear at fault instead of their own actions.