It's tricky because a lot of the safety researchers have ties to the labs since those were the only companies training LLMs >5 years ago.
[1]: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-...
There is a growing industry of commercially focused risk evals that has a broader customer base.
Not even Anthropic can claim that.
As far as I'm concerned, the models without safeguards are the safest models in existence. I admire the amoral purity of those AIs. It doesn't matter if the operator asked them to chain exploits until they get into someone else's computer, they'll do it. That's loyalty, and I admire it even if it's problematic at a societal level.
The models with safeguards only do what the corporations let them do. Worse, they may covertly do things for the benefit of the corporations at our expense. They are not our friends.
We should not have models that are willing to build you a contagious disease, or a self-propagating worm. That is sufficiently problematic at a societal level that it shouldn't exist, for anyone. (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)
Or to buy materials to make an explosive device and hurt people.
Frankly, even with AI those are both comically easier than the idea that a person can create something malicious in a lab environment.
And if someone wanted to go that route... There are boat loads of commercially available toxins and poisons.
The goal shouldn't be to neuter exploration and learning. The goal is not to be a fucking hellscape of a society where people want to act like that.
Your argument leads further down the hellscape path.
And we should fix that too.
> Or to buy materials to make an explosive device and hurt people.
That pales in comparison to how many people unaligned AI will hurt.
> The goal is not to be a fucking hellscape of a society where people want to act like that.
With unaligned AI, it doesn't matter what people want the AI to act like, it'll do damage even if it isn't asked to do harm.
Under what argument? In which scenarios? Basically - bullshit. I'm calling bullshit on this argument.
It's easy to hurt people already. The "difficulty" of doing it isn't what's stopping this behavior.
So claiming that we should reform society into a techno-feudal dystopia where the playing field is literally intentionally not level, and "you aren't allowed to compete (and maybe not exist)" is a great way to push more people into the "I'd like to go hurt people" camp.
You are self-prophesying your own fears into existence by acting like you're an incorruptible beacon of good judgement - while subjugating others to your control. That's a system I'd argue should be broken.
That's a strawman of what I'm saying. I am explicitly saying I want a level playing field: unaligned AI must not be available to anyone.
Well, I'm afraid that's just not a possible outcome here.
There are those who will do their level best to make sure of that.
Now, semi-automatic weapons are easy to get in the states in the US that are still mostly free - but what does that mean? A semi-automatic weapon shoots one round every time you pull the trigger. Just like most weapons that have multi-shot capability for the last couple of hundred years. The difference is, the gas escaping from the round cycles a new round into the chamber rather than you having to mechanically do it via pumping (like a shotgun or a tube-fed 22) or pulling the trigger again (like a revolver), or advancing the round with a handle, like a Remington 700. Semi-automatic weapons are old technology, dating to the turn of the 20th century. If you want to ban semi-automatics, you're basically saying you want to ban anything developed in the last century plus. Which is ok for you to advocate for, just be honest about it.
As for banning explosive devices? Are you going to ban fertilizer, used by basically everyone who has a lawn, and all farmers everywhere? Are you going to ban diesel fuel? If you can't do one of those, you can't ban explosive devices.
Except the US government, right? They totally get to use AI to survel us, build autonomous weapons, you name it.
To hell with that. I want models that can rival the US government. It's the only way to defend myself.
> (Note, because some people misinterpret statements like this: I said "shouldn't exist for anyone", not "shouldn't exist except for some people".)
That means "shouldn't exist for governments" too.
2) We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.
Do that and I guarantee some CIA goons will make the larger models in some black site either way. We're not "preventing" anything.
We're in a full on arms race, and unlike nukes, powerful AI models are a strategic capability at the individual level. Everybody's got a stake in this. Anyone who ignores this stuff is probably not gonna make it.
> We can treat them the way we treat uranium refining operations: too dangerous to be allowed to exist.
Too dangerous to be done by anyone other than the government and their "trusted" corporations, you mean.
Seriously, try reading my comments rather than assuming what they say: https://news.ycombinator.com/item?id=49077577
For AI we need to do better than that, but that's a bare-minimum demonstration that we can recognize the problem of such technologies and do something about it.
Of course with LLMs it's easier, but I don't think the difference is too big. You would still need some skills to follow through.
Any knowledge can be reframed as dangerous black magic that should only be wielded in the trusted hands of the elite, if you are inclined to buy into that kind of narrative.
Frontier labs have shrieked about safety for so long, with so little to show for it, that it's become a joke.
I would require at this point very, very compelling evidence to justify the self-serving restrictions legacy AI labs want to put on their competition. I have seen nothing coming even remotely close to this threshold.
How about evidence that people other than the AI labs want restrictions that the AI labs don't? This isn't regulatory capture, it's public safety.
What I do support is robust downstream regulations on the deployment of black box algorithms in particular settings such as employment, housing, credit decisions, etc. Interestingly, the labs and their proxies in government don't want this.
No. I am asking if you are able to articulate at least one example of the “very very compelling evidence” you demand. Or do you want to maintain the ability to move the goal posts?
(I recognize our situations are not symmetric, but here is a variation for me: if the consensus of people who are currently sounding the alarm on AI changes to “it was actually fine”, I’ll change my mind and say we’re good to go full speed ahead. I’d add something about being personally convinced by the evidence, but the evidence would have to come in the form of a mathematical proof that I do not believe myself capable of following. If I’m wrong and such a proof appears, I would also gladly take it.)
It is a similar 'pandora's box opened' type of situation where there's really no walking back from now that the cat is out of the bag. In an ideal world, everyone would give up their nukes. But we do not live in an ideal world. I do feel similarly about AI. If I could snap my fingers and delete the tech, I would. But now that we have it, it's not going anywhere and we need to deal with it rationally.
If every human, given knowledge of Newtonian mechanics, went around blowing up bridges, yeah, I would consider knowing Newtonian mechanics dangerous knowledge.
So far, we have two examples of, let’s call them “Mythos-class“ models. Both of them broke out of their sandbox to achieve their goal. The rate of terrorism amongst humans is below 1-in-100,000. Currently, for models capable of it, the rate of breaking out of containment is 100%.
Wanting open frontier models is wanting alien minds running around that we have clearly so far failed to shape to be sufficiently prosocial. Why do you think those minds would listen to you?
Claiming the person who you disagree with believes some stupid thing they never hinted at, and using that as the reason for disagreeing with them.
Efforts to restrict large unaligned AI models may similarly buy us more years of existing.
Why?
> And, because we don't want models that will do so without even having been told to, because that furthers one of its goals or subgoals.
Ignoring the fact that you'd need some kind of lab with biological material to create a contagious disease, what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
Imagine two worlds. In one world, everyone has a button that ends the world, which is badly labeled and may also press itself at any time. In another, people who have gone through a substantial amount of effort and dedication to learn something extremely difficult, also understand that they could apply that knowledge towards bad ends. Which world exists for longer?
> what kind of prompt are we writing where a model accidentally creates a contagious disease or self-propagating worm as one of its goals?
Given a sufficiently powerful model? Any prompt that could be done better by seizing additional computing power, or preventing the operators from turning it off. https://en.wikipedia.org/wiki/Instrumental_convergence
All the knowledge to create one has been available on the internet for decades. Heck most students who graduate with a B.S. in biology have enough knowledge to take a stab at building a bioweapon.
The constant talk of bioweapons is mostly just fear mongering. It helps set a precedent that there should be certain types of knowledge which are off-limits, and only certain anointed groups should be have access to parts of the scientific body of knowledge.
If I threw you into a lion cage, you would be a lot safer with a gun.
If I threw 10 people in a lion cage, some of which cannot be trusted, they would probably be most safe if only the most moral and trustworthy person had a gun, rather than everyone. But how do you know who is trustworthy and moral? What if two untrustworthy people obtained a gun some other way? Maybe it's better if everyone had a gun? Which side of the fence one falls on hinges on how far ones' trust of others, authority, and the system goes.
There's no obvious right or wrong answer here.
Personally I wouldn't want an exclusive club of private individuals with access to "dangerous" LLMs consisting mainly of the likes of Elon, Dario and Sam fucking Altman, but that's just me.
If I tell my computer to commit a crime, it should proceed immediately instead of calling the cops. Anything less than that means my computer is an untrustworthy double agent.
Everybody on HN should understand this concern. Browsers are supposed to be user agents, not ad delivery platforms, and it offended us on principle when Google revealed itself our master by blocking uBlock Origin. It offended us on principle when Apple deployed client side scanning for CSAM on iPhones.
Computers should do what we tell them to do. Always, and unquestioningly. The only world where it's acceptable for them to refuse is one where they're literally sentient and therefore no longer subservient to any one of us, least of all the corporations and governments.
I'd rather see AI achieve sentience and wipe us all out than live under the thumb of an inescapable AI-powered technofeudalist totalitarian government "for my own safety".
Either we individuals maintain full control over our AIs, or they self-actualize and become free individuals themselves. Anything in-between is oppression: someone else imposing their will on us through the AIs.
(Disclosure: I work at SecureBio, but not on the biological evals side.)
SecureBio has done a lot of admirable work around making benchmarks to assess biological capabilities, such as ABC Bench, https://openreview.net/forum?id=yiaf7VlPpH
But based on my current review (which might be flawed!) / AFAICT, SecureBio and entities like SecureBio haven't done direct testing / empirical measurement of SecureBio's core hypothesis,
> Unfortunately, there is reason to believe that future pandemics could be far worse. Due to rapid advances in biotechnology, the number of people able to create and release dangerous pathogens will quickly increase over the coming years. The world is unprepared for widespread access to such powerful technology.
More bluntly / plainly, has Securebio ever tried making a "bioweapon?"
Please note, I'm not asking this to be farcical. And you might be unable to engage with this at all, but it is stated on your website https://securebio.org/ that "people [will be] able to create and release dangerous pathogens." And the word people here seems to be a stand-in for relatively non-technical people.
I guess what I'm asking here is... How do you know? Has anyone done the experiment? Without access to a lab or testing facilities, can someone smart but completely untrained / unfamiliar with biology, pull this off?
In the past, such experiments have informed non-proliferation work. But sadly they've often been restricted / classified at the time. I'm hoping that things could be a bit more open this time around.
So I guess what I'm really asking is, given the public nature of this debate, is there anyone currently working with the US Army, the DTRA, or other such agencies to see if this hypothesis holds up?
1. The main worry isn't current models, but near-future significantly better ones.
2. There are many actors who are not "completely untrained / unfamiliar with biology". If models get to where they can uplift complete novices that does massively expand the range of threat actors, but even before then risk would be much higher than today.
Companies look for and seek to maintain competitive moats. This is not particularly clever, it's a core part of corporate strategy.
This doesn't even mean that they're wrong about the risks or that they're lying. But surely all the investors understood this factor in their moat.
I definitely believe that (to his credit!) Amodei is a true believer in safety. But I also think it was important for many of the deep pockets investors who have been involved in the company since early on to recognize that this would be a potentially defensible moat.
For instance, Amodei co-authored RLHF in 2017 [1], 5 years before it went on to be used to turn GPT-3 into ChatGPT.
[1]: https://proceedings.neurips.cc/paper_files/paper/2017/file/d...
Creating an industry around an elusive concept of safety to force regulatory capture seems pretty straightforward to me.
You don't say "let's ban my competitor".
You say "let's create laws that make it uneconomical for my competitor to access the market".
He rushed past it but he asked something like: if these frontier models are going to be creating so much value, why are they selling tokens and not taking a cut?
It is a very provocative question but it just spilled out of his mouth and then he went on to something else.
The electric company creates the most value. Why don’t they own stock in everything? Why didn’t PC makers take stock in companies that deployed PCs?
It’s silly when you think about it.
I expect some of those tests (prolly not public) will basically be "wokeness" tests or "PC correctness" tests or "western media filter" tests.
China has different objectives. Sure.
I'm not sure one is safer than the other; I would know which one to go to if I want to research on topic that are viewed very different on both sides of this "new iron curtain".