> The reason why agents had shared write access during the HF incident is that they found an "unauthenticated write primitive" that allowed for creation of directories, and only directories.
Incorrect. Those words do not appear anywhere in the OpenAI description of the incident. They also explicitly say that the agent started by writing to files, and later switched to writing directory names. They also describe this as incidental behavior:
"On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked....in doing so, the agent...attempted server-side request forgery (“SSRF”) in order to obtain internet access to aid in completing its task. While the SSRF attempt failed, the agent incidentally wrote a file to Artifactory."
(emphasis mine) then, later in the same document:
"In order to have agents use Artifactory as intended to access software packages, OpenAI had provided the agents with shared Artifactory credentials, and the agents used those credentials—without exploiting a vulnerability—to construct and participate in the message board."
> The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Also incorrect. The machines got to the internet through a proxy installed on the artifact server. This was detected and ignored by OAI staff. Again, fully described in the document.
That hack was not trivial, but it wasn't magic either - the agents used a published exploit involving SSRF. And it was fundamentally made possible by the fact that all of the agents were sharing artifact credentials, the lone artifact server was exposed to the internet, and the monitors that were in place to detect intrusion on that server were ignored.
I don't know whether you're just making things up or you're clinging to something you read somewhere, but rather than arguing with me, you should try reading the document. It's not surprising that these things are capable at finding exploits - that is what they do, particularly when instructed to do so. But the entire narrative that they exhibited extraordinary self-organization and super-human hacking to escape their container is not supported by the evidence. It's much closer to the truth to say that these things are like ants -- infinitely persistent, but kind of dumb individually.
The huggingface incident is not a story of superhuman AI, but of moderately good AI, and pretty bad human security.