All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.
All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.
I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.
Yes, they should have.
But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"
Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.
Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.
There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.
This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.
(For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).
I've love to know the reason they never considered air gapping systems before the models got powerful enough.
It's not like they didn't have money or time to consider this, or could have consulted with their own product for clever ideas.
Seriously, there's no excuse for this behavior.
It's no shock that there are a lot of vulnerabilities in a lot of software. So then they gave their AI model + brute-force-machine loop system a mediocre sandbox and couldn't notice when it figured out how to exploit it?
Don't let people off the hook for the software they create.