Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.
This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.
> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want
I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)
> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it
Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.
I'm challenging the notion that a model escaping a jail made by its creators, who are financially incentivised to make jailbreaking models, is meaningful towards the idea that the model is going to break out of a jail in the wild and do significant harm.
The examples being given by folks here, e.g. a model wiping an un-backed up home directory, simply doesn't strike me as being a unique problem in computing.
I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.
Not really. If I build a special new wine bottle, and call every breakage a mis-alignment problem, it's not the bottle just being fucked in the same way every fucked bottle is fucked, that's marketing. It doesn't change the fundamental form of the problem.
> "I am sorry your family is dead, my bad"
This should be punished. It's a problem that plagues Instagram and OpenAI. It's not inherently one, though, that has to do with AI. Just sociopaths preying on children.
> honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate
Perhaps. I haven't seen someone explain it to me in this thread in a way that seems separate from bugs.
Where I have seen a separate class of problem argued is where it's existential. But in that case, clarity of definition comes at the cost of any evidence for it.
You seem to be using a different definition of alignment from everyone else. Seems like it would be much easier for everyone if you just adopt everyone else's definition, rather than trying to convince everyone else to adopt yours.
You're still failing to provide the definition.
You're also falsely claiming your secret definition is universal. In this thread, someone claims deleting a home directory is a failure of aligment.
Robert Miles YT channel is a good place to start as it explains these concepts.
It's not limited to cyber attacks. LLMs helped terrorists learn how to jump motorcycles to assault a military base!
https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-...
the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.
It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?
I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
> In this case I don't think the intent of the user is to have the model break the evaluator
If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.
Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.
Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?
Is being aligned with the user prompt always a good thing?
I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?
Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"
Why do we think they're doing this? Nobody airgapped anything. Nobody pulled any products. We got a PR blurb.
Altman is a notoriour liar. Why would you give him the benefit of doubt? Based on the evidence, there is nothing here except a shrinking advantage over open-weight competition. Desperate men are shrieking for survival.
This is textbook misalignment. Literally the paperclip scenario.
2. Hugging Face did report this incident to law enforcement. (https://huggingface.co/blog/security-incident-july-2026)
3. If I hire a pentester, and in order to find a vulnerability they hack into a third party that has some information about my systems, the pentester has done something wrong. If I ask a model to solve a CTF challenge, and it goes out and hacks Hugging Face to find the answers, the model has done something wrong. I think it's fair to call this kind of wrongdoing misalignment.