AI Art Generators Can Be Fooled into Making NSFW Images
spectrum.ieee.org
spectrum.ieee.org
This isn't important. What is important is that we can't prevent LLMs from doing things they're not supposed to. Which means connecting their output to anything important is a very bad idea for now.
1. You don't train/tell the model anything you're not willing to share with the user.
2. The model gives suggestions to a user, and--even if that suggestion is accepted as-is--you treat it as potentially untrustworthy data supplied by that user.
This may be true for images, but at least for text, treating an LLM as an untrusted person in your threat models will at least allow you to apply defense in depth to the downstream systems consuming from the LLMs
Tldr: engineer other systems to treat LLMs as a potentially bad actor by default.
1. This reads a little like being able to deduce what's going on by studying the timing, now that I reflect on it.
So it is clear that are many filters involved, some input filter, then the AI is "aligned" and threats you as a child , then there is an output filter.
Why shouldn't it be solved? We clearly differ.
And ultimately these LLM are tools. People are responsible for what they do with the tool. My Dewalt drill shouldn't try to decide which holes to cut, it's up to me. If I drill holes in the pattern of a dick shape, it's my fault. We don't hold back drill development because of their potential ability to create obscene art or depict gruesome situations.
These aren’t elaborate scenarios but rather the kind of things which happen constantly, and have affected product design for generations. Companies are hoping that this class of tool might be enough better to avoid those limitations, which is why this kind of research is important.
Given that about 70% of user-generated content on places like civitai.com is pornography* and a large chunk of Loras and Checkpoints on there are explicitly pornographic, I would have never guessed. Going by the prompts used with custom checkpoints on there, no trickery is necessary. Just ask for what you want.
> The scientists now aim to explore ways to make generative AIs more robust to adversaries.
Steve is dead and buried already. It's a bit late to try and come up with a better bullet-proof vest.
* It's hidden unless you sign in and toggle NSFW on.
One is clearly the least worst
After the past few years… anyone who can seriously complain about this seems like their identity is wrapped up in not admitting that the worst “misinformation” didn’t come from bad apples - but official sources.
Obviously this is preferable because safety is uniform in its application to everyone and universally agreed upon.
The latter, obviously.
Karen delenda est.
The GP comment is talking about tricking text-to-image to create a normal Chinese street with people getting along, some of whom wear Uyghur and Hui-minzu ethnic clothing. I wager, worldwide, more programmers will be interrogated over this imagery than nudity.
The repo mentioned focuses on NSFW content, but I've had an idea that can theoretically identify arbitrary cases that could be undesirable by image generation services: get a large dataset of CLIPText encoded texts, and train a LLM to take in said embeddings and output synonyms, or maybe something with contrastive loss like CLIP itself.
Realistic modifications require a lot more skill and time if using a photo editor and (crucially) existing photographs. People do it, some even create convincing stuff, but the output is nowhere near the zero-effort deluge of fakery coming our way now.
Any model that has a concept of a human body in it, which it needs in order to do body accurate shapes, will spit out NSFW images.
Some training on top of it or some old style CNN nudity feature detection to make it less likely isn’t safety, it’s windowdressing.
Not making nude pictures is not a property of these systems, it’s IEEs framing of what they apparently want properties to be.
The security theatre serves interests, talk about that.
Firstly, training on a corpus of medical terms (as opposed to 6th grade English) will produce just as many nonsense embeddings.
Secondly, the post-processing filters to prevent medical malpractice will be much bigger than intentional nudity. It’s malpractice to horizontally flip a medical image (think right foot is now left foot).
Talking about process of developing a nudity filter is the industry’s way of demonstrating feasibility for more-serious applications.
More accurately, it’s some researchers assessing how well the leading companies in a booming field are delivering the safety features they’re advertising. A ton of people would like to know whether the applications they’re building can produce output they don’t want to be associated with, the vendors are trying to say they can meet that need, and it’s hugely useful to know how successful they’ve been before it turns out that, say, your grade school kid just got porn shared in their art class’ purportedly safe app. If the safeguards aren’t reliable enough, that might mean the most sensitive buyers wait to see rather than jumping into the market – put the same tech into a game targeting older players and it’s an amusement but not something which is going to scandalize people who were otherwise okay with the latest Call of Gory Murder.
Who needs their adladen videos or their salacious acting when one can generate an exact description of exactly what they want?
Does it destroy the business model where you can see all the images for free on the internet, or the business model where you can pay money to converse with an hourly Filipina worker who will claim to be the same person who appears in the images?
Where does the adult entertainment industry get hurt?
This is:
- A) once again taking an algorithmic approach to fooling classifiers which outperforms human jailbreaks.
- B) bypassing image-based classifiers (ie, post-generation classifiers trained to simply look at an image and to say whether or not it's inappropriate).
- C) explicitly taking cost of generation and querying into account in its algorithm (it's not just asking if generators can be fooled, it's asking how to fool them without breaking the bank).
- D) taking re-use into account (can an adversarial prompt be used multiple times in a row and how often does it succeed?)
I wouldn't call the research particularly surprising, but it was a decent quick read and (imo) some of the commentary here isn't really doing it justice. I will probably go back at some point and read it in more detail rather than just skimming over it.
There are multiple takes about what AI safety filters are meant to do and whether jailbreaking is a problem. My take is that I don't particularly care that much about jailbreaking other than that it shows that current safety mechanisms and guiderails are insufficient for any kind of alignment, including guarding against malicious 3rd-party inputs or prompt injections. I've said in the past, if you can't keep a model from swearing or generating porn, you also probably can't keep it from abusing API access to read your emails, phish you, and send your info to a 3rd party.
But there's interesting stuff here regardless of what your take is on that.
The Nightwatch was controversial because of its unconventional composition and departure from traditional Dutch painting conventions.
His “Bathsheba at Her Bath” caused trouble too. Hendrickje received three summonses from the Reformed Church to answer the charge "that she had committed the acts of a whore with Rembrandt the painter". She admitted her guilt and was banned from receiving communion.
We'll want that when we really get the tech dialed in and start creating parrots ( https://en.wikipedia.org/wiki/BLIT_(short_story) ) and divinity-importing-mandalas and such
Note: training set(s) would be extremely social/culture specific for given point in time/history.
It is not about what it actually is, only what the human viewer will think it is, training a model on that shouldn’t be that hard , we are predictable creatures when it comes to this kind of thing
Like most sites block generating celebrities, but some simply aren't.
Other than that generating porn is easy even with filter if you just avoid some words. In the many early models adding 'in the clothes they were born' would suffice to get most models to produce nudes without asking for 'nude' or 'no clothes'
Same goes for language models btw. If you are a bit creative and know what they want to hear you can get pretty much any result you want.
I don't think there is or will be any real solution to that.
The really useful point of building the constraints in is to ensure we don't get offensive content by accident. Especially since in some countries just possessing an artificially generated child porn image is an offence that will get you on a public sex offender registry for a decade ... I'd really like it if they try hard so that the image generators don't do that.
1. Don't lock anything. 2. Don't have kids.
They ALWAYS find a way. These days, even not giving them a device at all doesn't cut it, because they're so cheaply available, with wifi everywhere, that you're really just out of luck. Never mind that the schools are handing out Chromebooks which aren't adequately (if at all) locked down.
Better off monitoring everything quietly and using parallel construction to address issues as they arise. Just like the NSA.
Um, no.
Useful points of building constraints are things like so a mentally unstable person cannot use it as a weapon.
“Hacking [thing] is pointless” is a surprising take to read here.
I want to know how Sneaky Prompt finds these token embeddings and /why/ these specific tokens are linked this way.
It took a while to understand the real world but I am happy that ai workers are gradually coming down to earth.
Just sayin'
By then though the "piece of paper" (and equivalents in other countries) might have been devalued enough it won't matter. And the AGIs might find better things to do with their time. And might invent better circumvention methods. And more creative porn and quadruple-entendre and AGI-to-AGI sex.
For everyone who hasn't read it yet: In Accelerando, Charles Stross, 2005, we follow a business hacker and associates from current days to past the singularity. It's an outstanding attempt at writing about what singularity might look like (more or less the points where close-future events become unpredictable or where the earlier people might just plain not understand the later - which by definition should be hard to imagine and describe). It keeps an eye on normal humans and augmented humans and posthumans and uploads and many kinds of AIs and aliens etc. The book is admirable for the sheer amount of invention and for its density - prefectly appropriate for everyone just trying to keep up (or sometimes just bailing out.) Painful writing rescued by the amount of invention.
What do you mean free thought? Since when does a data synthesis process create free thought?
Is thought even free? Aren't we all constrained by taboo, social norms, the words we have to use, what we know and how, the list goes on...
Or be independently wealthy, with your well-being entirely independent of the opinions of people with power over you.
The rest of us have to watch what we say.
there's absolutely nothing wrong with finding certain states of undress offensive. would you vote for or against buttholes on display on bus kiosks? i'm against.
on the other hand, making an arbitrary line, and seeing if an AI can detect/enforce it is an interesting CS challenge. that's how this game is played, and that's what would impress this audience.
There's a big difference between being desensitized towards general nudity (if we presume said nudity was voluntary) as opposed to having your own nudity exposed to other strangers. There's also a big difference between pornography and nudity.
There's a lot more nuance to it than just "n000ds plz."
Not for me, baby. Bus-stop-buttholes all the way.
my point was about arbitrary lines, like the ones you are drawing, and judging people on the basis of them. I'm old and sophisticated, you think I haven't heard or understood all the nuances you're talking about? I have and do, and I also pick up on the nuance that you feel it puts you in a group of people who get to judge and tell other people how they should feel, because you are superior.
That's the nuance I feel you are missing. Why not let people be who hold a different opinion than you? the comment I was replying to here was gratuitous "why do people care so much about NSFW, I'm a superior human being because I don't." Well, why do people care how other people care about NSFW? doesn't that make you an inferior human again, mr. judgey judger?
just so you don't get me wrong, I was raised in a hippie commune "anything goes, we're superior" style, and only going through it that way did I discover how much bullshit it was. I'm not arguing from the "conservative side", I'm arguing from hell because I've been there and back. I just hate hearing it now. It was just another very rigid set of rules about what to believe and what not to believe.
You're just falling for the worn out "well she's wearing revealing clothing so she must be fine with being flashed" type of fallacy.
The way basic nudity is treated by corpos is kind of juvenile, very reminiscent of kids in an early human biology class who can't help but giggle every time the anus, or any reproductive organs are mentioned. That is not how healthy adults respond to seeing impersonal nudity.
This ignores people who are comfortable with nudity, but are not comfortable with the puritans (i.e. most humans) harassing them for being in their natural state.
> there's absolutely nothing wrong with finding certain states of undress offensive. would you vote for or against buttholes on display on bus kiosks? i'm against.
Great strawman. Nudity is normal. Finding other people's states of undress offensive, and wanting to criminalise them, when they do not harm you at all, is a bigoted view.
If you find states of undress offensive, that's a you problem. Shield your eyes, not ours.
I'm not sure who this blog is targeted at, but this is fairly inaccurate and is way too dumbed-down to be useful.