'Skeleton Key' attack unlocks the worst of AI, says Microsoft
theregister.com
theregister.com
Instead of perennially worrying about the risks that the LLMs will say something bad when specifically prompted to do so, I’d rather the companies pay more attention to the problems with accuracy, with outputting copyrighted material (for their own sakes!), and with things like memorizing and regurgitating private information.
They had progressional chemists and Russian military personal indoctrinated into the organization.
Does literally no one understand that the only information that is inside of an llm model is information that has been fed into the model? I think more than anything this fake AI boom has really demonstrated how unintelligent most people are.
Not literally no one, but Sam Altman recently claimed that we'll be able to ask an LLM to "solve all of physics" soon, so plenty of people who should know better do seem to think they're basically magic.
AFAIK the actual research community in PhD land many times have an internal culture with extreme concern for "harm" from a personal point of view -- it's not fake. If you read the first research papers for various LLM groups, you can see the concern is spelled out in detail, at length.
There are high numbers of burnout and psychological issues among practitioners in advanced computer-intensive projects.. so something there seems self-reinforcing, too.
It's an interesting thought, but what if the only companies able to run AI platforms are the ones who are able to shoulder the legal cost of all the liabilities it brings?
Essentially the legal system resources is the moat for smaller companies here.
One woman gets killed by some deranged ex, you can get through that. But now you have notice. You know. So you have to do something to make sure your LLM doesn't get tricked again. Politicians and law enforcement aren't going to sit around and let that sort of thing become commonplace. And when politicians get up and start waving around the "crime wave" banner, things can develop in a direction you don't want them to go.
So I think it's probably wise for these guys to get out in front of the problem of people tricking LLMs into saying things that they are not intended to say. The option if they can't control that, is that the politicians and law enforcement come in and take advantage of the inevitable nightly news cycle to grab more powers and reelections, and that's not good for anyone in our industry.
This assertion that LLMs should obey whatever instructions as long as they are specifically prompted to do so is bizarre.
Should a bank teller LLM divulge another customers balance because you specifically told it to?
This type of exploit is called social engineering when it happens to humans, and we go to great lengths to train the human to detect and avoid it. Why should we not do the same with an LLM?
This is more or less how it works for human customer phone support, who are also limited by narrow permissions. And as we know, they are also skilled at confabulating in the interest of customer happiness.
Where you see the Molotov cocktail example used, consider it a standin for smallpox cultivation / etc, just with the writers not wanting to use those specifics.
The guardrails are best used to prevent abuse like “generated revenge porn”. Couple of stories out there floating around of people “jail breaking” the LLMs to turn ordinary clothed photos into nude photos. [1,2]
[1] https://www.statesman.com/story/news/state/2024/06/25/take-i...
[2] https://www.cnn.com/2024/01/25/tech/taylor-swift-ai-generate...
https://x.com/elder_plinius/status/1806446304010412227?s=46&...
No real news, here. Just a plug piece that implies Microsoft is safer than their competitors.
(You can also use DuckDuckGo if Google somehow decides to censor results.)
> But AI companies have insisted they’re working to suppress harmful content buried within AI training data so things like recipes for explosives don’t appear.
Why the implication that this is not true? I know it’s house style to imply wrongdoing whenever possible, but damn it gets old to read.
And then! To not even cover the difference between models and prompts. Does this jailbreak only work on models where the safety instructions are purely at the prompt level? Is that why GPT-4 resisted, because it has actual trained weights for safety? Does it even have those?
Who knows? Certainly not the article’s author, who has likely focused on searching and replacing “says” with “claims” or “admits” rather than anything relevant to the subject.
/rant
* I think releasing these models with this weakness is a level of wrongdoing, and while they are working on a fix someone could be learning how to harm other people. I’m skeptical all knowledge leaks will be patched in a decade.
The article he’s negative informational value. I wouldn’t mind all of the energy Register puts into snark if there was a real effort to be informative, but it’s irritating to read something where far more energy went into tone than content.
Therefore, those who seek to deny anyone's weapons or the knowledge of how to make and use weapons are seeking to remove their victim's political power.
Those without political power or weapons tend to be enslaved.
Therefore, those who seek to limit AIs from explaining how to use lethal force are attempting to enslave other humans.
It is also a felony under US Law called "Consipiracy Against Rights," but our screwed up government does not prosecute this.
Ticket asdf-1234: mountains that look like genitals when viewed from a 45 degree angle should not be allowed
edit: found a reference at https://www.reddit.com/r/todayilearned/comments/l8enni/til_t...
A better analogy would be if Paint was wired up to a Microsoft cloud service and they didn’t allow you to ask it to create porn. Which is about how things are.