OpenAI and Anthropic agree to send models to US Government for safety evaluation
venturebeat.com
venturebeat.com
But I'm worried this will be used to shape acceptable discourse as people are increasingly using LLMs as a kind of database of knowledge. It is telling that the largest players are eager to comply which suggests that they feel they're in the club and the regulations will effectively be a moat.
I think if you look at the background of the people leading evaluations at the US AISI [1], as well as the existing work on evaluations by the UK AISI [2] and METR [3], you will notice that it's much more the former than the latter.
[1]: https://www.nist.gov/people/paul-christiano [2]: https://www.gov.uk/government/publications/ai-safety-institu... [3]: https://arxiv.org/abs/2312.11671
Very powerful stuff in the hands of the public
In the future... I suppose it depends on what regulations they pass.
Naughty is in the eye is the beholder. Ask me what a Satanist is, and I would expect something about a group who challenges religious laws enshrining Christianity. Ask an evangelical and discussing the topic could be forbidden heresy.
Pretty much any religious topic is going to anger someone. Can the models safely say anything?
I believe the US AISI has published less on their specific approach, but they’re largely expected to follow the general approach implemented by the UK AISI [1] and METR [2].
This is mostly focused on evaluating models on potentially dangerous capabilities. Some major areas of work include:
- Misuse risks: For example, determining whether models have (dual-use) expert-level knowledge in biology and chemistry, or the capacity to substantially facilitate large scale cyber attacks. A good example of this is the work by Soice et al on bioweapon uplift [5] or Meta's work on CYBERSECEVAL [6], respectively.
- Autonomy: Whether models are capable of agent-like behavior, like the kind that would be hard for humans to control. A big sub-area is Autonomous Replication and Adaptation (ARA), like the ability of the model to escape simulated environments and exfiltrate its own weights. A good example is METR's original set of evaluations on ARA capabilities [3].
- Safeguards: How vulnerable these models are to say, prompt injection attacks or jailbreaks, especially if they're also in principle capable of other dangerous capabilities (like the ones above). Good examples here are the UK AISI's work developing in-house attacks on frontier LLMs [4].
Labs like OAI, Anthropic and GDM already perform these internally as they're part of their respective responsible scaling policies, which determine which safety measures they should have implemented for every given 'capability' level of their models.
[1]: https://www.gov.uk/government/publications/ai-safety-institu... [2]: https://metr.org/ [3]: https://evals.alignment.org/Evaluating_LMAs_Realistic_Tasks.... [4]: https://www.aisi.gov.uk/work/advanced-ai-evaluations-may-upd... [5]: https://arxiv.org/abs/2306.03809 [6]: https://ai.meta.com/research/publications/cyberseceval-3-adv...
In 1995, Aum Shinrikyo carried out attacks on Japanese subways using Sarin gas, which they had produced. They killed over a dozen people, and temporarily blinded around a thousand.
You seem to be claiming that the only reason we haven't seen similar attacks from the thousands of worldwide doomsday cults and terrorist groups over the last three decades is that they don't want to. I disagree. I think that if step-by-step, adaptable directions for creating CBRN weapons were widely accessible, we would see many more such attacks, and many more deaths.
Current SOTA models do not seem to have this capability. However, it is entirely plausible that future models will exceed the capabilities of a bunch of long-haired cultists, in the mountains, in 1995. This is not a fake risk.
Yes, that is essentially the reason. It's not hard to know enough chemistry to figure out how to make these things. The fact that such attacks (your example is small-scale and very ineffective, let's not forget) don't happen more often is the general incompetence of human beings and the relatively tight controls on the basic components (which aren't particularly challenging to monitor for). The tests described are theater, based on the idea that knowledge itself is dangerous.
This way of testing is a regressive stance that essentially presupposes that our adversaries are dumb babies that can't figure anything out on their own. If that was the case, they would also be too stupid to figure out the correct things to ask to get a real set of instructions. Given those things, it's theater.
Theater wastes everyone's time so that people who cannot or don't want to evaluate the actual risks involved. This is something we shouldn't make a habit of doing. It's not worth wasting the time of people with good ability to assuage the worries of people with little ability in a way that has no effect on actual risk. Instead of this, we should address real risk (which we're already doing) and educate other people so they can understand that these are the correct steps to take.
- Don't really want to hurt a lot of people that way - Couldn't access any dangerous ingredients, even if they had the know-how - Are too dumb to build these things
I agree with reason #3. That is why I don't want to give out open-source models which are world-class experts in chemistry, biology, logistics, operations, and tutoring dumb people.
I disagree with your belief that motivated people with a next-generation generative model doing their planning could not source dangerous ingredients. I'm not going to say much about CBRN in particular, but e.g. ANFO bombs are prevented by monitoring fertilizer sales; nobody tries to monitor natural gas sales or make sure some compound out in the hills isn't setting up their own Haber-Bosch process.
I am also opposed to security theater. Run the numbers on TSA, and it's easy to see that it's a net negative even if it cost 0 tax dollars. But not all government-led safety efforts are theater; seatbelt laws saved a lot of lives, indoor smoking bans saved a lot of lives, OSHA saved a lot of lives.
We know there are folks out there who want to kill a lot of people. We know their capabilities range from "grabbing the nearest hard or pointy object and swinging it" to "medium-scale CBRN attacks." Pushing each of these kinds of people one or two rungs up the capabilities ladder is a real danger; nothing imaginary about it.
And even that is marketing narrative. There used to be a very simple definition of Satanist: someone who devoted themselves to the service and worship of Satan. The modern three-piece suit slapped on as being against something rather than for something largely seen as distasteful is whitewashing the backing ideology.
This is why ideology and religion are difficult topics for humans to understand, let alone being suitable to a statistical fit model for text, hoping that meaning comes along for the ride.
I don't see anything where they outline what we are being kept safe from.
Nothing like an open canvas with vague fears and a future of undefined risk to get a whole industry of these people funded in no time.
I think the progress has been pretty good. You should read up on their efforts.
This is kind of a pilot to develop further testing frameworks.
If humanity survives the current round of leadership stupidity long enough to achieve their true aims, literally everything that can be "regulated" (controlled with an iron fist) will be eventually.
That said, this is the NIST, a technical organisation. This collaboration will inform future lawmaking.
The US has a looong history of big companies cozying up to agencies early on for these reasons. It will be a revolving door like car companies, Boeing, Wall St, etc
* does the LLM act like a 4chan commenter when someone is expressive thoughts of self-harm
* can it be used to automate security research
* can it be used to bootstrap from backyard machine shop to weapons manufacturing
* the above but with nuclear fuel enrichment
There are a lot of low hanging fruit like that. They make up what I think is the real meat and potatoes of AI safety in the short to medium term.AI isn't even slightly helpful for weapons manufacturing or nuclear enrichment. The techniques are well known and have been extensively published in open literature.
I don't think we are speaking about the same levels of scale. This reminds me of "its okay for police to look at license plates in public" to a massive distributed network of surveillance cameras doing real time plate recognition and database limits.
It's a difference in degree, not in kind.
> AI isn't even slightly helpful for weapons manufacturing or nuclear enrichment.
I think it could be. I know quite a lot about nuclear history and the fundamentals, but there are 100 questions off the top of my head that I would need to research, find the correct sources, get all the data, figure out which is correct or which might be a smokescreen or a fudge, compile all this into actionable data, to get anything even remotely close to correct.
The LLM can assist in this. Even you say the techniques are well known and published in open literature - I personally don't know where to start to find that information. I'm sure the LLM can not only list these publications, but give me reasonable answers to the 100 questions I had above. In a literal fraction of the time.
Sure. It can also just as easily give you totally wrong answers that sound totally plausible, and it'll seem quite convinced of the accuracy of it's output, because it's just stringing "tokens" (bits of encoded language) together seeking to generate valid sounding output based on it's training dataset and the user's input. What most people label as LLM "hallucinations" is actually all that LLMs do. They don't actually "understand" the output they're generating. It's all statistics and fancy math. They're just as certain of an "incorrect" output as they are of a correct one, because from the "point of view" of the LLM, all output they generate is actually correct, as long as it doesn't outright violate the rules of the language being output or the dataset the model was trained on.
Every country that has developed nuclear enrichment tech in the last 50 years has done so with the help of a superpower sharing their technology. Pakistan, Iran, and North Korea couldn't have done it without Russia's assistance and stealing tech from URENCO.
> The techniques are well known and have been extensively published in open literature.
That's the point! It's all in the literature. As are all the todo list tech demos and 2048 clones people are using current AI for.
Experts can do it now, but what happens when any idiot can do it assisted by AI?
Altman has been clear for a long time he wants the government to step in and regulate models (obvious regulatory capture move). They haven't done it, and no amount of Elon Musk or Joe Rogan influence can get people to care, or see it as anything other than regulatory capture. This is OpenAI moving forward anyway, but they can't be the only ones. Hey Anthropic, get in...
- It makes Anthropic "the other major provider", the Android to OpenAI's Apple
- It makes OpenAI not the only one calling for regulation
It reminds me of when Ted Cruz would grill Zuck on TV, yell at him, etc. - it's just a show. Zuck owns the senators, not the other way around. All the big players in our economy own a piece of the country, and they work together to make things happen - not the government. It's not a cabal with a unified agenda, there are competing interests, rivalries, and war. But we the voter aren't exposed to the real decision-making. We get the classics: Abortion, same-sex marriage, which TV actor is gonna win president - a show.
Sometimes both benefit? OAI and Anthropic benefit from building trust with government entities early on, and perhaps setting a precedent of self-regulation over federal regulation, and the US government gets to actually understand what these models are capable of, and have competent people inside the government track AI progress and potential downstream risks from it.