Search 5.8B images used to train popular AI art models
haveibeentrained.com
haveibeentrained.com
Can't they crowd-source a proper labeling project - I wonder how much better things like Stable Diffusion would be if its training would include correct, complete labels for the images. I'm sure lots of folks would willingly spend a few minutes here and there to aid with the labeling if it means they get to enjoy the model for free.
Imagen and Stable-Diffusion both used subsets of this full 5.8B image set.
> While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset which is known to contain a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes.
Other mediums don't have enough people interested in curating and labelling 5 million images with very detailed tag lists for free.
Pretty sure they are trying to figure out a way their art can’t be used to train an AI and won’t be willingly participating in making it easier.
We don't see this as binary in the long term. Maybe artists want to release art from a prior period, for example, but withhold their current series until they move to the next.
Who told you you could use the data in the first place?
It's copyrighted. Reproducing any of it is a crime.
In which case it should be fairly simple to challenge it in court, no?
This is not true of the artist market. While yes, many artists get into it for the fun/satisfaction of producing their own artworks, the reason that they can remain in it long term is that there are other entities in paying for the results of their work. See: game studios, film studios, etc. Before AI, the only way to get the assets for these various projects produced was by paying artists/experts for their work. Now, the production of those building blocks can be automated.
The LAION-5B set in particular is filtered by it's CLIP embeddings - basically run a better AI to classify the images, and check that the two descriptions rate highly enough for similarity, which in turn lets the system have some confidence an image isn't completely wrong (also the step you filter out NSFW and other things if you want, generally).
So for instance for photography generation what may be needed is huge amount of clean photos with obsessively detailed labels, maybe just the exact same single-point-lighting (the exact coordinates of the light being a data point, plus strength/lumens), with 8 pictures of each subject in black background: front, back, left, right, top, bottom, 3/4 mostly-front (AKA the corner), 3/4 mostly-back, and then the same 8 ones but with white background, then also include all the info possible: weight, height, width and depth, plus versions of the most common states of each object (ball: inflated or deflated; bird: flying, idle or walking), plus photos adding two subjects together (one dataset of woman with hat, another wearing the same clothes but without the hat, and one of just the hat without the woman), with properly labeled relationships so it's clear they all refer to the exact same hat (or lack of), you get the idea...
Then the most important thing for "image generation AI" may not be computing power but the most boredom-resiliant staff you can hire; of course that's just an example for photography, for things like text you may need an equivalent rigorous effort by a multitude of linguists.
They could use Wikimedia Commons to get a smaller collection of better labeled images. Currently it's images from Common Crawl extracted by Laion, not sure if it already includes Wikimedia Commons.
The ESP game[1] paired two random people looking at the same image while a timer ticked down. The players had to enter labels that described the image, and if both players applied the same label, their score increased, and the label was associated more strongly with that image.
Well-known labels for an image were excluded after a while, so you had to guess less and less obvious labels in order to score as time went by. Once in a while the system would also assign test images with known labels to prevent cheating.
Apparently a lot of people played it because it was fun, so there wasn't even the need to pay them for labeling the dataset...
I'm very surprised how primitive it is. I put in a few names of uncommon technical objects and some of its images were close and others so far off as to be unrecognizable or totally unrelated/useless.
What it needs is a quick/instant feedback system that allows humans to rate an image against the query word. A rating scale of say 1 to 5 where 1 is an identical match, 2 a close match, 3 marginal, 4 ambiguous, 5 wrong/totally unrelated. If in place then likely thousands would make an effort to optimize the matching.
" Releasing our first Spawning tool to help artists see if they are present in popular AI Art training data, and register to use our tools to opt in and opt out of AI training
I think we have created a way to make this work out well for everyone"
and StableDiffusion's Emad Mostaque agreeing to support the initiative https://twitter.com/EMostaque/status/1570158985852121090
It's built off of common crawl, so it probably does have a pretty representative sample from whatever the big image searches use.
Funny enough, the NSFW filter that laion built is turned on. Without it, it's... a lot. The NSFW stuff is done with a model, so you get a probability of NSFW out of it, and you can select a threshold.
If you set the threshold high, like 95% certainty that an image is nsfw to filter it, you get a bunch of false negatives, letting a ton of nsfw through. Set it too low, and you throw out stuff that isn't nsfw.
We (haveibeentrained) erred on the side of too high, so we wouldn't tell artists their work wasn't in there if it was. Tough trade-off there. Similar to using the dataset to train an AI model, where you might cut off useful images from training if you try and filter all the nsfw.
"Rutkowski" returns a bunch of book covers, repeated a lot. Can you ensure images returned have diverse embeddings? I expected digital art, not detective stories.
Do you use CLIP or just metadata?
What is intended process that starts after you get artist's email (for either of purposes).
In the next few weeks we'll be adding the ability to log in and flag or upload your works (if they aren't there). Those lists will have permissions assigned to them, starting with simple opt-in or opt-out.
The goal here is to give people the opportunity to remove images they don't want in this dataset or add images they do want in there.
When we enable sign-in, we'll also add a privacy policy, because at that point, we will store some images, on request, to use them to make finding other works by the same artists easier.
Opt-out image URL lists will be made available to the dataset owners for removal. Opt-in image lists will be public.
For example, type in "Filipino" and you get NSFW photos.
Filipino is what we call citizens of the Philippines, btw.
That's not a joke.
If Spawning is able to have your images removed from the training set by version 1.7 or whatever, you will be removed from any models actually in use for real commercial applications.
All rights reserved. I didn't agree for them to use my IP commercially.
EDIT: Let's consider a simpler scenario - I take an MD5 sum of one of your photos, and then hash collide a an image from my phone camera till I find a match. Which part of this process is "stealing"?
Seeing as you like analogies: if you play music in a public building, you need a license even if you are selling coffee and not the music.
It's the same as if I were to look at your photos and then take my own similar photos for the background of a web store. I would have "used" your photos in some sense. But I would not have used them commercially in any way whatsoever.
You have just as much grounds for complaint in Stability's case as you would in the case of me looking at your photos as reference for my own.
Likewise there is no copyright infringement occurring when an ML model is trained on copyrighted works.
Stable Diffusion is a 5GB matrix of floating-point numbers trained on 240TB of data. It does not, and cannot, contain infringing data from the training set. It is physically impossible for it to contain such data. There is not, and cannot be, any infringement occurring.
That's called "non-literal" copyright infringement. Plaintiffs win on the claims and appeals courts uphold the judgments. (See, e.g., the "Blurred Lines" case. https://www.nbcnews.com/pop-culture/music/robin-thicke-pharr...).
A person could infringe the copyright in an original image if they painted their own version by hand. The AI is basically irrelevant.
So there's almost a kilobyte of missing data, without which nothing is produced.
Would a JPEG file with random noise in it be potentially any given original image? Of course not - even though the decompressor is perfectly capable of recreating one given suitable input data.
The "Grokking Stable Diffusion" colab posted here a week or two back makes this particularly explicit - any given 512x512 image can be reversed into a latent-space encoding of that image, and reconstructed from it. The NN weights are necessary but not sufficient to do so - ultimately you end up with a 512 byte number mapping to the expression which reconstructs an image. But that includes images which aren't part of the original training set.
[1] https://colab.research.google.com/drive/1dlgggNa5Mz8sEAGU0wF...
Slam dunk case right? What's stopping you?
The part they’re concerned about is where somebody uses their work as a reference for derivative works.
We can acknowledge that this is an unsettled legal question without pretending like there’s no such question for a creator to raise. This community works better when we give each other fair understanding.
That's not the standard for copyright infringement. The standard is (1) access to the original work and (2) producing a work that is "substantial similar" to the original. If the user of an AI produces an output that is substantially similar to one of the AI's training images, then it could potentially infringe.
That's the law. In an actual case, a judge or jury would look at the original and decide if the alleged infringer is "substantially similar" enough. That's a crap shoot, so cases usually settle before that.
So the question “where in the model is your original?” is a reasonable and relevant question to ask.
If this person can induce the model to produce a work that is reasonably considered a copy of their original, then fair enough. All they have to do is give a prompt and a seed and they can prove copyright infringement very easily because anybody else with similar hardware and software can demonstrate the infringement on demand. If I understand correctly, this has happened with GitHub Copilot, with Copilot reproducing copyrighted works verbatim.
But if they can’t do that… why should anybody take their claims of copyright infringement seriously? If nobody can point to copyright infringement having taken place, what basis is there for believing it has? As you say, the standard is access to the work and producing a work that is substantially similar to the other. The former has been demonstrated. People are asking about the latter.
So “where can we find the original in the model?” is possibly the single most relevant question there is. It’s a clear line that divides infringement from inspiration, and it can be proved definitively if copyright infringement has been observed to occur.
A copy must be made for there to be copyright infringement. Who has shown a copy has been made? If a copy has been observed to occur, it’s trivial to demonstrate copyright infringement. Who has done this?
---
Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following:
(1) to reproduce the copyrighted work in copies or phonorecords;
(2) to prepare derivative works based upon the copyrighted work;
(3) to distribute copies or phonorecords of the copyrighted work to the public by sale or other transfer of ownership, or by rental, lease, or lending;
(4) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and motion pictures and other audiovisual works, to perform the copyrighted work publicly;
(5) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and pictorial, graphic, or sculptural works, including the individual images of a motion picture or other audiovisual work, to display the copyrighted work publicly; and
(6) in the case of sound recordings, to perform the copyrighted work publicly by means of a digital audio transmission.
----
This is far from an unlimited grant of power. Of this list, the only plausible grounds is (2) - derivative works.
Which means we're well into then arguing about "Fair Use"[1], which I would encourage people to read the full description of carefully - because the answer isn't whether you can come up with a snippy "gotcha!" it's whether under careful consideration in the court of law anyone would be likely to agree with you.
I really don't know why people bother bringing it up other than to create an uninformed distraction.
https://people.com/music/pharell-williams-robin-thicke-blurr...
I capture what I see, more as a frame of reference for myself with location scouting and game dev. I take quite literally hundreds of thousands of photographs. It's 1 in every thousand that's worth sharing.
I have zero interest in other photographers, I can't even name one. Whatever other people do, neat. I don't even consider myself to be a "photographer", but I guess with hotels and airports buying prints, that somehow makes me a professional.
I capture what I like. There's literally zero other motivation behind it. Most of my income comes from elsewhere. I use a print-on-demand service.
If I take a photograph of a sculpture I've made and that is then incorporated into another work without permission, that's intellectual property theft. A lot of the images are of things I've made, not just street photography or landscapes that anybody could go and take.
If I write a riff in a microtonal scale and it turns up as a sample in a record, or is covered and released without permission, that's intellectual property theft too.
If someone is worried about an original image, why not watermark all uploaded versions?
If I walk about my neighborhood scattering paper copies of an image, and someone else picks one up to hang on their wall, did anyone break any laws (besides me littering)?
More bullshit equating the human mind to an AI. Newsflash, boyo - humans are not NNs, there is no reason legal precedent should treat humans the same as NNs (and it won't).
It's the law that matters here.
So - keeping this argument purely in terms of IP law - can you explain what the case is for these images to have been illegally used?
Sort of falls flat.
Related: we recently released[2] semantic search for custom datasets. You can use it to find all sorts of weird stuff in benchmark datasets like MS COCO[3] used to train many computer vision models.
If folks are interested I can write up a "how we made it" post describing the behind the scenes.
Once uploaded, it's Stolen!
{username}'s profile image
We'd love to have you sign up if you're interested! We expect to start the beta for those tools in 2 to 3 weeks.
I haven’t tried but I don’t see SD making any memes by themselves yet.
https://newsletters.theatlantic.com/galaxy-brain/6317de90bcb...
https://www.inputmag.com/culture/mat-dryhurst-holly-herndon-...
And when they mix the two things get interesting. Having topless women selling (cow’s) milk was a very strange experience as a “prudish” American.
Yes, and the guy who took these "porno chic" pictures was outed as a creep and harasser.
No, that's why it's called pornography, very much to differentiate it from something "noble" like art, it's vulgar and trivial, by the very definition of the word it isn't art.
"Noble" is nowhere in any working definition of art. Some academic snob might use it with their in-crowd. Damien Hirsch's shark parts is considered art.
"everything I say is art is art", you basically. Well anything I say isn't art isn't art. Especially pornography in general, which sole purpose is sexual gratification of the audience through sexual exploitation, rape via human trafficking.
There is nothing snob is saying that. There something egregious and dehumanizing in saying what you say on the other hand and trying to normalize the very definition of vulgarity.
Not necessarily; plenty of pornography contains serious artistic contents and merit beyond the simple aim of sexual gratification, and they raise interesting questions on their own. This is why some have attempted to make another classifier of 'erotica', which IMO isn't needed.
>through sexual exploitation, rape via human trafficking.
Even if this were true for all porn (and it isn't), that dosen't make it any less art; it may make it 'lesser art' from a moralist perspecitve on the value of art, but this by no means of widely accepted relevance to aesthetic value nor to the property of being 'art'.
>the very definition of vulgarity.
A rude gesture or crass words seems vulgar. A video of a penis being sucked doesn't. "Vulgar" is the sort of word you might hear a shrill, paternal sitcom character use to describe a minor inconvenience or faux pas, and it's just as hilarious when people use it as an argument in real life.
In the mean time you can spare people gratuitous, graphic descriptions of explicit sex acts, AKA pornography, and keep your depraved fantasies for yourself.
Your certainly qualify as a vulgar individual given the fact that you feel the need to depict pornography here trying to make a point about god knows what.
Point about god knows what? Pretty sure it's been clear all along, art encompasses things we like and things we don't.
to depict: to represent or characterize in words; describe.
https://www.dictionary.com/browse/depict
Then you keep babbling about your life like anybody cares.
If you think a blowjob being described in seven words to make a point about vulgarity is 'gratuitous', 'graphic', or 'explicit', your moral compass is severely out of whack.
Only some people define porn as degrading and rape-ful. Galleries are full of paintings of slaughter, is art about snuffing out life? Is all erotica porn? Where do you draw that line?
You sound anti-sex in dismissing sexual stimulation as non-artistic. I could stimulate rage and sadness in you with a painting of two girls stabbing a kitten in the eyes, why is that painting less art than a painting of two women making love to each other? Sex is a good thing, even outside of marriage imho.
Pornography isn't sex, it's a gross caricature of sex, usually shot from the perspective of male subjects involving the degradation of women, other men or children, purely for commercial purposes in order to give a false sense of sexual gratification. This isn't art, this is akin to butchery, with human bodies as meat ready to be consumed by incels and all kind of other frustrated males and depraved individual.
The former certainly like claiming pornography has any artistic merit as a way to justify their consumption of depraved content and their lack of any actual sex life.
Sex with another person is difficult to come by for many, and only comes sporadically for many more. You don't sound like you've been married for 20 years and bound by oath to not have sex with anyone other than someone who's physically lost all libido. I recommend some empathy.
You keep on making assumptions about your interlocutor, stick to the matter discussed.
I'd recommend you stop trying making things personal and have some empathy yourself for the victims of the porn trade and sex trafficking instead of trying defend pornography by labeling its critics "anti-sex". Porn is not sex and whoever gets off watching porn isn't having sex at first place. The fact that people have a non-existing libido is the least of my concerns.
You talk about sexual stimulation for people that can't have sex with other humans as though it's a bad thing. That's just being a prude.
You conflate sex with porn, then try to paint porn sick individuals as "victim" because they are not entitled to have sex, while being completely oblivious to the real victims of porn production and trade, trying to paint pornography as any other thing than violence, grotesque and vulgar and you have the audacity to call me prude because I'm against sexual violence and its effects on society? If that's being "prude" in the mind of porn sick individuals trying to paint their fetish as art, then I'll be "prude" all day.
Things can be "artistic" or contain art within them, but it doesn't make them art.
I don't buy that something done/created for the purpose of being viewed and evoking (even strong or many) reactions is enough of a criteria.
Otherwise things from trolling to 9/11 are "performance art" (an already barely hanging-on category), the latter of which was very "successful" at meeting this minimal set of rules.
The corpus is of art, not Art.
On the plus side, i appear to have not (yet) been trained. I can consider myself safe from the AI. For now.
We are building an opt-in list, because a lot of people do want to be able to prompt AI with something like, "a cat in the style of me" or "me riding a dinosaur". That will be shared publicly, of course.
You can't really back an image out of a model, you can only retrain without that image.
Creative Commons should add a new license that prohibits the use of ones work to train AI.
{"message":"Forbidden"}
in the api response and the page says- Sorry, there was an error with your search. Please try a different request.Why isn't this immediately being used as a missing persons (or pet) database service?
1) Privacy.
2) Lack of geolocation data associated.
3) Privacy.
4) Lack of contact information attached.
5) Privacy.
6) Lack of case numbers for various missing persons cases being attached.
7) Privacy.
I was thinking more that 1,2,3,5,7 were all privacy related.
I also have fib burned into my head from pivotal tracker story points...