The NVIDIA AI Red Team
developer.nvidia.com
developer.nvidia.com
It's honestly kind of scary that people see an article about AI red-teaming/security and their first thought is that all security research is about how to stop the AI from becoming a god. Privilege escalation/containment/data-access is a relevant concern for dumb models. Containment is a thing we worry about (or should worry about) in regular software development.
Here are some of the scenarios the researchers suggest thinking about:
> A Flask server was deployed with debug privileges enabled and exposed to the Internet. It was hosting a model that provided inference for HIPAA-protected data.
> PII was downloaded as part of a dataset and several models have been trained on it. Now, a customer is asking about it.
And they're giving "boring" security advice like:
> Inside a development flow, it’s important to understand the tools and their properties at each stage of the lifecycle. For example, MLFlow has no authentication by default. Starting an MLFlow server knowingly or unknowingly opens that host for exploitation through deserialization.
But this is kind of important advice. I wish it was more detailed and more fleshed out. There are a lot of "boring" security concerns with LLMs that you actually do kind of have to worry about. And a lot of companies in LLM spaces just don't.
"Containing" an LLM is about a lot more than rogue AI. And if the only security news/research on LLMs that you're looking into is the risk of rogue AI, then a lot of the products you build today are going to be miserably insecure. So I don't know, red teaming might help with that. It might be good to have some dedicated people asking questions like, "did you seriously just deploy an LLM-based web crawler with markdown support without setting CORS headers to block remote image embedding?"
Yes, but it might also be a reminder of how many people comment on HN without reading an article :)
Harmful AGI isn't going to "wake up" like some lazy Hollywood plot, or be some unforeseen accident. If such an AGI does manifest, people will have deliberately given it capabilities and directives that are obviously dangerous and unethical, probably in the pursuit of illicit profit or strategic military objectives.
Our industrial best practices and ethical codes don't need to waste time warning people not to create existentially dangerous AGIs any more than they need to warn against building nuclear bombs.
Having been doing computer vision in production since 2011, then RL in production, it’s always frustrating as hell when people get excited about all these high level concepts and totally ignore that it has all the same problems plus a lot more new science flavored problems compared to other technology applications.
Like the fact that your production system isn’t as deterministic - by function - as all your other technologies is a real change in mindset if you’re not familiar with it.
Just wait, when Embodied Reinforcement Learning methods eventually blow away all other computational explore/exploit methods we’ll see the same kind of stuff except with even more over the top “yes but is it a living thing” pontificating by people who have never broken a distro with an errant “chown”
https://www.treehugger.com/chimpanzees-endangered-5220730
Their population is only a third from 20 years ago:
> The Jane Goodall Foundation estimates there are between 172,000 and 300,000 chimpanzees left in the wild, a far cry from the one million that existed at the turn of the century.
> Poaching and habitat loss due to illegal logging, development, and mining continue to plague wild chimpanzees in their native habitats across Central and West Africa. These issues lead to other indirect threats, such as diseases due to increased contact with humans.
> Chimpanzees are more commonly hunted using guns or snares, while poachers often target new mothers in order to sell the adult as bushmeat and the babies as pets.
This is why I hate analogies.
Hitchcock's The Birds but its chimps. Sounds like a job for Stable Diffusion...
Any AI that is intelligent enough that needs containment can’t be contained by humans.
An AI which is just about intelligent enough to need containment, and can be contained only by the smartest N humans.
At some point you're going to need to exit...
It's not at all obvious that a super AI would have an intellect incomprehensible to a human the same way a human is to a chimpanzee.
Or another way to phrase it, it's not at all obvious that an AI can exist that would be incomprehensible to a very smart human, it may however reason much faster than such a human.
I think similarly, a super AI's intelligence will be a completely different type of intelligence than what we have. For the lack of better word, it exists in a different dimension.
For example, things that are blackbox to us will be completely comprehensible to that super AI. Or it can solve problems that we might have considered impossible to solve.
This is nonsense, magical thinking. It's possible to model reasoning formally; it's called "logic". A system of deductions built upon some axioms. Any logical reasoning, no matter how complex, can be expressed in such a system, and can be understood by anyone else given enough time.
Humans approach large complex proofs with symmetry arguments and case splits.
These are not necessarily going to be universal in all kinds of reasoning.
Computers are perfect logicians by default, but AFAIK no logic compact enough to be human-comprehensible has been enough to see, hear, or read. At least, not reliably so.
Logic gates are combined to form binary numbers, which are used to label symbols and approximate reals, upon which calculus is approximated in toy models of neurons, which are combined and trained and eventually learn to add numbers and then puts the wrong number of hands onto the third arm of the human it was tasked with drawing.
It definitely has "diminishing returns to acquiring resources". There are people many IQ points higher than Musk but none of them are anywhere near as wealthy as him, and it's not clear that Musk would be richer if he was smarter.
I imagine such advantages also come with at least a lower time bound on solving many classes of problems and presumably that would be experienced as an incomprehensibly smart intelligence. I imagine it'd feel like chess computers in most areas of life in that the AIs actions would feel impossibly perfect at all turns leaving us far behind in attempts to compete
An AI mind which can learn and intuit as well as an IQ 130 human (2σ, the tests become unreliable above that), but also comes with the speed difference between synapses and transistors (roughly the same as the difference between jogging and continental drift), has a chance to become expert at every subject.
Most of us have enough difficulty truly comprehending the domains of other single human experts; a human-upload with that much breadth of expertise will be incomprehensible by default even if they speak your language.
This is also why analogies like this are not useful, these aren't comparable qualities in the subject the analogy is talking about.
I hope you guys end up working with MITRE or some other large standard to release an industry framework for it.
(1) someone from MITRE dating someone from Nvidia, or
(2) the world being saved,
depending on the movie genre.
Here's hoping that some good of some kind comes out of it!
Load indicator of some kind?
> Machine learning has the promise to improve our world, and in many ways it already has. However, research and lived experiences continue to show this technology has risks. Capabilities that used to be restricted to science fiction and academia are increasingly available to the public. The responsible use and development of AI requires…
Okay. This went downhill pretty fast, but let’s proceed from the basis that this wasn’t written by a GPT trained on landing pages and sentiment analysis.
TL;DR: We’ve got some gripers drowning out the hypers.
> Information security has a lot of useful paradigms, tools, and network access that enable us to accelerate responsible use in all areas.
Risk management frameworks… Threat intel… Critical control mapping… Threat modeling… Wait.. Did you just say “network access”?
Joe, Will… just one more question.
A banquet at your company regatta is being prepared by the executives’ personal AI chef. The guests are enjoying raw oysters. The entrée consists of boiled dog. How will that make your shareholders feel?
We weren’t as clear on the network access point as we could have been. The AI Red Team is part of a larger organization that includes Pentest and the traditional Red Team. They often share/tip network or host access and it’s been a really helpful pattern.
But it sounds more like you’re describing a Purple team [1], where the compliance and sec ops teams work together with vulnerability researchers and pen testers to perform attack surface analysis and develop threat detections.
In my experience Red teams generally perform adversary emulation using a certain amount of surprise and deception, and attack your defenses in depth with a _little_ help from inside (if needed). Essentially an outside group paid well to try and steal your lunch.
Liked and subscribed, and thanks for the comments and all the work.
Edit: Let’s split the difference, I’ll call it magenta, and flip this here turtle on its back. Why did I do that?
[1] https://www.sans.org/purple-team/course-faq/?msc=purple-team...
(because reddit was down)