LLMs already have a body count.
Adding safety controls on LLMs makes about as much sense as adding safety controls on TempleOS because the random messages are getting too prophetic. It's as if all the leaders and captains of industry have devolved into some primitive, weak-scifi shamanism.
This whole "discussion" about "AI safety" is about giving them more runway to avoid delivering quantifiable value to investors for a little longer while they "figure things out." The great consensus from the valley is that everyone needs internal (and therefore bullshit) controls. Trust us now! But nothing with real teeth that would require a costly regulatory and compliance framework.
Nuclear at least is supposed to be air-gapped, in practice this has been imperfect.
As demonstrated with HuggingFace, such AI driven hacks can be a surprise even to the people who instructed the AI, both by happening at all and also because they can targeted at entities who are not even truly relevant to the instructions given.
The most obvious failure mode for their hacking evals was an improperly configured, tested and monitored sandbox.
Similarly, the very first question after an impressively correct result from any ML tool, LLM or not, is to see if the answer was already in the training data.
These companies don't even handle the blatantly obvious failure modes that do not kill people.
If I leave a gun in my front closet, it won’t independently walk out the door and go shoot people, no matter what I might say to it — unlike an LLM.
The gun does not. No matter what, I have to choose to pick it up, aim it at someone, and pull the trigger.
Where exactly is the flaw in the analogy?
In my model evals, if a model deviates from the provided prompt (task adherence), that’s a failure even if the primary goal might have been achieved in a different manner.
To go with the analogy, task deviation should be treated the same as a gun that due to manufacturing defects can fire despite the safety being on. That defect remains, even if you can use the gun to shoot (in an unsafe manner).
Simply, neither should happen and both models deviating from their prompt or guns firing by themselves are to be considered a fatal flaw. It's why, despite greatly lauding the GPT-5 series, which did adhere to prompts in most every scenario, I have ranked every OpenAI model post Spud very poorly as those traded task adherence for brute forced, deviated approaches to solutions and why HF, Medicare, etc. were inevitable with their current trajectory.
A model that as part of normal, well scoped use proceeds by taking independent action or, far worse, makes choices beyond the original prompt, is not something I feel should be used. GPT-5.6 Sol and GPT-6 Astra both do this on the regular when trying to safe an ancient, utterly messed up git tree with branches upon branches that I maintain for eval purposes, thus leading to data loss that if the models adhered to the prompt as written, wouldn't happen (though the task will take three times more steps). Something GLM-5.3 Flash, prior OpenAI models including original GPT-5, anything from Anthropic in the recent years, etc. do not fail at.
Nothing happens without an initial prompt, deviating from it is a severe flaw and should lead to a model not being considered for deployment or wide use.
Yes, alignment is important. No, alignment is not perfect.
And at the end of the day guns don’t kill people, people kill people. But a model isn’t like a gun.
Do I think it’s likely? No. Do I think AI doomerism is a joke? Yes. Does that change anything I’ve been saying? No.
They are both tools, ideally properly implemented by the manufacturers and (improper implementation/defects not withstanding) require active operation by a user before something happens.
To go back to the original example, don't interact with either an LLM or gun in the way they were designed to be used and neither will do anything.
If you want to keep that analogy, use a voice-activated gun. Doesn't make it any less of a tool, doesn't make it any less of the users sole responsibility, doesn't mean misinterpreting the users inputs or just acting blindly when a user asks to "protect me" isn't a failure that should have the product taken off the market. Even more so if that happens in a "sandbox" as part of "safety testing" that "accidentally used impossible to solve tasks".
If a prompt wasn't clear enough, the model must ask for clarification and pause over using brute force.
I couldn’t come up with an argument as tortured and ridiculous as yours if I tried.
Liability and safety requirements, when needed, should be placed on final product manufacturers, not the tools they use to build things, whether pencils or LLMs. My 2c.
:(
There's some angles from which this distinction doesn't matter too much, because begging bad actors not to develop a similar harness won't work. But the frontier labs seem to be taking it for granted that the LLM is the only part of this system that matters, and we don't need to ask any questions about whether their commercial products should ship with a harness that's allowed to execute unvetted code and spawn hundreds of subagents.
https://en.wikipedia.org/wiki/Anthropic–United_States_Depart...
That said, when the problem is at the level of "the government itself is breaking the law", you can reasonably ask if any regulation is even worth the paper it's written on.
What you want at this point, given the government lust for it, looks more like a bunch of countires saying ~"we consider development of autonomous weapons[0] by to be a casus belli and will go to war to prevent it, and also that development of same by private individuals anywhere in the world regardless of normal sovreign territorial limitations[1] is equivalent to acts of piracy on the high seas".
[0] But then you'd need a more precise definition of "autonomous weapons" to avoid accidentally including a Phalanx CIWS etc.: https://en.wikipedia.org/wiki/Phalanx_CIWS
[1] So much for Westphalian sovereignty :/
Yeah, exactly, and ultimately I think that's really the thrust of the point I was making.
And, to me, if I was just looking at this calmly as a decision about what the obvious direction seems to be, given these factors, it's pretty straightforward: deprecate the nation-states. They are the ones mucking up the whole system.
If the thing we're really concerned about is LLM-safety wrt warfare and weapons, then I'd much rather tell the (whining, childish, seemingly headed for self-destruction anyway) nation-states that they have to sit this next era of humanity out than have to nerf them for the rest of us (and as you point out, nerf them in a way that the nation-states won't abide anyway).