92 karma · joined May 26, 2026
Also, I listed the view points I had to show that it seems other more reasonable first assumptions don’t seem likely, therefore, the last potential of this being either faked or carefully not avoided seems more likely than the others (based on the current information we have).
Could you clarify which point or assumption you are objecting to?
Still, to address your comment about what you’d expect to see in world 2 and 3 (assume 1 was true), that’s why 1 was addressed separately. I don’t believe I argued that the potential for world 2 or 3 prevented world 1.
As for the ‘evil exec’s strategy’, I would call this a mild incident but if it were much less I wouldn’t guess they would get a lot of press. The press coverage is certainly repaying the token cost as well. If it was planned, it seems to be going well given the press coverage I’ve seen on it. So I wouldn’t assume the plan lacked enough to weaken the idea that it’s a plan. But to be clear, my stance is just based on the info I see now which isn’t a lot… subject to change.
As for the containment piece, if you were testing an AI model on its hacking capabilities that you believed was far more capable than anything you’ve seen, I would assume you would air gap it (a network control). Done right (no signals ability) I would argue this could be next to impossible to break out of. But it’s a fair jab to say I should have added some qualification on the “impossible” piece as next to nothing is truly impossible.
Your comment about jailbreaks being more one off and hard to do consistently in agents is a good point. Still getting an agent to hack isn’t hard even without a jailbreak, you just have to tell get creative in what you tell it. I’ve found telling it that it’s in a CTF or that I own the system that it’s hacking will work fine. A lot of offensive security companies are running agents in their testing so getting an agent to hack seems commonplace.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
There is a similar way to do this casting with infusion and sintering using metal powder infused FTP printer filament. It shrinks down a bit more than this freeze casting sintering I think but it skips the slurry/freezing step of freeze casting.
This actually isn’t true. Having done physical security work before, a weird fact is that one of the best physical deterrents is lighting; even over CCTV.
I don’t say that to take away from this, this is great work and I’d love to see the lighting toned down for multiple reasons. However, this should be framed as a security tradeoff not an outright win.
Source: The Impact and Policy Relevance of Street Lighting for Crime Prevention: A Systematic Review Based on a Half-Century of Evaluation Research (https://www.crimrxiv.com/pub/wl9zqxga/release/1)
The electric car is the new daily driver, the gas car is an old daily driver, but the stick is an old work truck that’s been around for probably over 40 years since its basic (you see pavement when you open the hood). Even if the old cars with sticks get replaced, construction equipment and machinery may always benefit from simple gas engines… even if that’s not ideal for the environment.
What I envisioned for how it works is fairly similar to this, QUIC can actually be more difficult to detect than it seems since it’s very dynamic.
I’m not saying you’re wrong, I’m saying you can’t tell from this incident.
I’m not saying I recommend LastPass for that reason, but I wouldn’t write them off for that reason.