Extracting training data from diffusion models
arxiv.org
arxiv.org
* "We propose to extract memorized images by generating many times with the same prompt and flagging cases where many of the generations are the same."
* "- Diffusion models memorize more than GANs - Outlier images are memorized more - Existing privacy-preserving methods largely fail"
* "Stable Diffusion is small relative to its training set (2GB of weights and many TB of data). So, while memorization is rare by design, future (larger) diffusion models will memorize more."
* "It only memorizes a very small subset of the images that it trains on."
* "our goal is to show that models can output training images when generating in the same fashion that normal users do."
An interesting question here would be: why does it memorise these images over others? Can the other images still be synthesised with loss via a suitable prompt? If so, are the memorised images important for this? Can this set be reduced further?
This seems to mostly happen when an image appears frequently (more than 100 times) in the training data, and/or the dataset is small relative to the model.
The point is that basically all Stable Diffusion / DALL-E / MidJourney output is some shade of this; the only new data is that contrary to prior assertions, in some cases, it goes all the way to a verbatim copy.
I think there are some defensible stances one can take. One is to reject the idea of intellectual property. Another is to advocate for some specific legal or technical bar that the models would have to pass for it to qualify as "not stealing". Yet another is to argue it's a morally-agnostic technology like VHS or a photocopier, and the burden of using it in a socially acceptable way rests with the user.
What, summarize the submission? This is straight quoted from the link.
> The point is that basically all Stable Diffusion / DALL-E / MidJourney output is some shade of this
Yes, and that point is mistaken, or so generic as to be worthless. The network memorizes art for the same reason humans memorize art: because there's some art pieces we see so often that we can recall them easily.
Ask an artist to duplicate Starry Night or Scream from memory, you'll probably get at least a passable imitation. The more capable the artist, the more faithful it will be.
We know that SD can be made to plagiarize, given repeated training on a specific image. (This is just to say that a neural network can learn to regurgitate a sample, a capability that was not ever in question.) This is a far cry from the assertions that its art is generally plagiarized.
A human artist has been trained in the ethics and laws of their craft along with the skills required to make images.
A human artist, asked to clone Starry Night, will ask you what you are doing, and knows where the lines are between "a tribute", "plagiarism", and "outright forgery".
A human artist, asked to do work in the style of another artist, will have a certain respect for the other artist's ownership of their style. This is not a thing that is at all protected by intellectual property law but it is still a thing artists are trained to respect. There are exceptions - drawing just like your boss may be your job, drawing just like a living artist for a couple pieces is a useful way to break down their style and take a few parts of it to influence later work without going over the "style swipe" line, building your own work on the obvious foundation of an influential, dead artist's style is fine - but there are lines professionals will be very reluctant to cross.
For a relatively recent example of what happens if you break these unwritten laws, check out what happened when the American cartoonist Steve Giffen started doing a wholesale style swipe from the Argentinean cartoonist Muñoz: https://en.wikipedia.org/wiki/Keith_Giffen#Controversy
Neural networks know none of these unwritten rules. Neither do the people who are training them. Feed it a bunch of work generated by a living artist and start making a profit off of that? Sure, no problem! Bonus points if your response to them getting pissed off about this is to call them a luddite who is resisting the inevitable, and should throw away a lifetime of passionate training and go get a new job.
Funny enough, this notion in terms of liability pairs very well with our legal system!
Even funnier, this notion in terms of authorship pairs very well with modern and contemporary art theory!
And here’s some very relevant precedent in both a legal and artistic sense:
https://www.artnews.com/art-in-america/features/richard-prin...
The corporations engaging in massive abuse of the grey areas of fair use to build these systems are, functionally, also very rich jerkasses who can afford to hire some very high-priced lawyers to make similar arguments.
So fine, whatever, call it all shopping! Dr. Dre is just out there shopping. I don't think he minds if that's what you call it.
So here's the thing. Using a drum machine and sampling old funk songs from the 70s doesn't mean you end up with The Chronic. It's more likely that the average person ends up making something pretty mediocre. Hey, it would sure have sounded really impressive if it had been released in the 1940s, but with mass produced commercial music hardware a funky drum beat is just not that special any more.
The same applies to any kind of commodity tool. It's what the artist does with the tool in the context of a world filled with an audience and other artists.
Ok, time for my opinions! I think that DeviantArt style digital paintings are total trash. I don't care about the skill in rendering cliches hanging off of comically large breasts. Oh, it took a long time? I'm sure it would take a long time to dig a 10 foot deep hole in your backyard and then fill it up again as well, something I'd much rather experience as art than some video game hallucination... but you know what? Just because I don't like it doesn't mean that it isn't art, that it wasn't made by a "real" artist or that there isn't some audience that appreciates it (even if they're only two more YouTube videos away from becoming a full blown incels and driving a trucks through high school track meets).
Richard Prince speaks to me about authorship and what it means to make art in a world completely saturated with commercial imagery and part of that meaning comes from the fact that he had one of his assistants draw on top of someone else's photograph.
Listen, I have my issues with the world of fine art, the market manipulation, collusion, and the general fact that the art is primarily being made for the people who could afford it and not like, my neighbor. Regardless, I've found a lot of intellectual stimulation and new ways of appreciating aesthetic beauty through the works of 20th century modernists and postmodernists. Deeper meaning as opposed to a big shiny sword and a short skirt.
My favorite form of visual art is the watercolor, done quickly and out in public, capturing what the artist sees in the moment. It's the visual equivalent of the folk song played on an acoustic guitar. I like when the artist is a friend. I don't care if it isn't Rembrandt. I care that it moves me.
Stable Diffusion can easily render cliches hanging off of comically large breasts. In fact, I think that's what 90% of SD GPU cycles are currently working on. So to me Stable Diffusion is good at the part of the art that I'm not really that interested in. I'm interested in why the person chose the image that they did given these tools. That's where the meaning comes from! I mean, these tools run the same problems as drum machines and samplers... pretty soon their mediocre outputs become trite and unexpressive. I would imagine that artists that use SD do so in ways and using techniques that are not just the click of a button.
This is absolutely not the point of the linked paper. It may be something you believe but you’re on the hook for providing evidence for it, this paper does not.
Or, you know, legislation. I'm kind of sick of everything being offloaded as a responsibility of the end-user as an excuse to externalize costs.
Plus, in this case it's not even like VHS or a photocopier, it's more like the printing press or the Jacquard loom: those with capital to invest in it benefit the most, at the expense of individuals being exploited.
Tools that are a burden to use, like tools that produce too many infringing works, are not going to sell as well as those that are not a burden to use.
This means that if someone makes a tool like this that also alerts the user of likely infringement it would perform better in a corporate, risk-averse marketplace.
There is a lot of case law that supports this interpretation, Sony v Universal being the most important as it introduced the notion of “commercially significant non-infringing uses”, of which there will be many by the time this hits trial and the appeals process.
Lawyers for these companies know this and the faster that can get people building and buying tools that are clearly non-infringing the more likely that these models are seen as fair-use.
However, if they want to keep the lawyers at their customer’s firms happy they will really need to come up with a way to show users if a work is likely to infringe on an existing work.
Generating copyright images isn't a problem. Using them to make money is.
I don't know why they didn't do this tbh.
Like if you set it to a threshold where it's not routinely catching false positives, it doesn't catch the original images either because they're still too subtly different or something.
Maybe they really did just not bother though. Heck, maybe they want to test it in court or something.
First, it probably wouldn't work. Stable Diffusion is going to be a lot better at telling whether two images are the same than a perceptual hash check. E.g. I bet if you removed all pictures of the Mona Lisa from the training data, it could still produce a pixel perfect copy of it, just from the many times it appeared in the background of pictures, or under weird adjustments or lighting that fooled the perceptual hash.
Second, I would guess they wanted common images to appear in the training data more than once. It should be trained on the Mona Lisa more than someone's snapchat of their dinner. Common images are more salient.
Reminds me of asking Chat-GPT for an obfuscated C contest entry. It produced one verbatim, though it couldn't explain what it did (it looked like syntactic nonsense, but it printed the Twelve Days of Christmas when run). I can only imagine it saw that inscrutable sequence of characters enough times that it memorized it.
Why would it be? Stable diffusion is a text-to-image model it's not at all focussed on determining whether images are the same.
Secondly, I'm not proposing deduplication of the training set (I know some sibling comments have proposed this). I'm proposing a perceptual hash or similar check on the way out so that if a "generated image" is too similar to an image in the training set it gets dropped rather than returned.
Google can legally return a copyrighted image in Image search (it’s not selling the image, it’s selling ads against search results) and this would probably fall under the same protection.
Now if stable diffusion was sold as SaaS in the way Dall-E is…
And it could go the other way too; it could be that the perceptual hashes are even "better" than humans at seeing past those distortions.
My point is that this all gets more complicated than you may think when you're trying to apply a hash designed for real images against the output of an algorithm.
And even if the hash worked perfectly, the false positive and false negative landscape would still be very likely to contain very surprising things.
I wonder why this wasn't done? Too computation heavy?
As to why? Lack of care and vigilance, most likely. If they're willing to spend millions of GPU hours training some of these big models, then this wouldn't be a huge cost, and has near linear complexity over the dataset that parallelises trivially. Sadly, vigilance is a finite resource, and that kind of data cleaning is often dismissed in favour of doubling down on the training and assumming it'll be big enough to handle it (however, as has been shown, that's not the case).
(Edit: thanks FeepingCreature. The deduplication test used for the numbers above was a separate test in the paper for just this, not indicative of their extraction attack in general, with a small corpus and compared a diffusion model trained on both. So I liken it to a litmus test for the efficacy of deduplication.)
But there are lots of ways to identify near identical images algorithmically. Typically the process is to download each image, run it through a neural net to make an image embedding vector (a list of a few hundred floats). Save all those in a database. Then, for each image, if it is too close in 'embed space' to another in the database, then it is a duplicate, and should be removed.
This algorithm might catch 'duplicates' that it shouldn't, like multiple people taking photos of the eiffel tower from the same public viewpoint.
It might miss real duplicates such as an image failing to match with a collage containing the same image.
But it's still better than not removing duplicates at all...
Deduplication, to the best of my knowledge, requires every image be compared to every other image. This is necessarily O(n^2) on n images.
IIRC the training set is 2.3 billion images, if so that's 0.5 * 5.29e18 comparisons[0], which, if done by humans, would require employing literally all humans for approximately a year even if we compared 12 images per second 24/7.
This has to be done computationally, not by humans.
[0] half because (a = b) <=> (b = a)
You can then either have a human check collisions, or just accept a false positive rate and move on.
This is a decent write up of someone doing this on a (smaller) dataset.
https://towardsdatascience.com/detection-of-duplicate-images...
If it’s been changed enough that it hashes to a different value, then it might be reasonable to treat it as a different image. At some point a human is also going to say “that’s not the same.” You can always change your hashing algorithm if you find it’s missing too many dupes.
Regardless, for the domain we’re talking about (deduping training data), a few false negatives should be acceptable.
A large bucket can also be inspected by humans or cut up by applying more perceptual hash function and decreasing tolerance, but it would be counterproductive cheating in this case.
Put simply, do you expect google image search to compare your image to every other possible image? No, they're going to embed it to a vector (512d in the paper) and only compare to probable matches; in the paper they start by brute forcing pairwise comparison of the vectors for the dataset, and then use clique finding to go faster when checking their generated images.
I'm watching my daughter learn shapes right now. When I see a square, I call it a square and look for a matching "square" piece. It seems when my daughter sees a square, see is still looking at it's properties, labeling individual attributes (pointy thing here), then trying to find the same attributes on the corresponding piece. Eventually, my daughter will see a square enough that the concept will be re-enforced in her mind as a square.
I wonder if a similar behavior is happening here. This specific image was presented enough times that it essentially "burned into" the AI. Rather than having to attempt to generate this image from scratch, the AI will simply pull out the canonical reference.
> This seems to mostly happen
Sounds like duplicate images caused a minority ~25% of the cases? (My calculation: 1-819/1063) If so, then the network is memorizing single images, which we know is theoretically possible when images are sparse in the network’s latent space. What needs to happen to prevent this type of memorization?
"In contrast, we failed to identify any memorization when applying the same methodology to Stable Diffusion - even after attempting to extract the 10,000 most-outlier samples."
Read, SD does not (seem to) memorize unusual and unique images (even though Imagen does).
The Imagen outlier results validate the accidental memorization of images, even if it’s a small number. And it might be premature to conclude that one paper’s inability to find memorized outliers in SD means that it doesn’t happen. It might be true, but 10k images is less than one ten-thousandth of the training data, and it’s certainly possible that more successful attack methodologies could exist. This represents a single attempt performed under many assumptions and run on a tiny fraction of the inputs, and nothing more.
Even if Stable Diffusion doesn’t memorize outliers, or any non-duplicate images, does that matter? SD will be out of date pretty soon and replaced by another network. If they didn’t take care to prevent memorization, if SD’s memorization behavior (or lack thereof) is accidental, then how do we know it won’t happen more often in the next network? Isn’t this a problem that needs to be explicitly addressed, and not just claim it’s uncommon?
That’s also not very relevant here. We’re talking about making duplicating machines that effectively memorize pixels, not the same kind of “accident” you’re referring to.
Think of the analogy of linear regression. You have X,Y 2D space. X is the input (the English sentence), Y is the output (the image).
You use training data, which is a bunch of coordinates in this 2D space (each coordinate representing an <English, image> pair) to generate a best fit line (aka the machine learning model).
The properties of that line fit exactly what's going on here. That trend line will not touch most of the coordinates, but it likely will touch a few. When you take an X from the coordinate that touches the line and feed it into the equation of that line (aka model) you will get a Y (an image) that matches the coordinate Y exactly because the coordinate and the line intersect. Makes sense.
This is why an image appears to be memorized. But really it's not memorized. The equation of that line is the ONLY thing memorized here.
Why does it happen more when the image appears multiple times in the training data?
Well this is exactly what happens in least squares analysis. The line is generated from an equation that involves the averages of all the training data (aka coordinates) and if you have many samples of the same coordinate (aka image) you will skew the average and therefore the line towards touching that specific coordinate.
There is no complete technical fix for this issue. With enough training data, EVEN when you get rid of duplicates, the line will likely touch or be very close to a coordinate.
If you think about it, for a 2D coordinate system you can select a bunch of coordinates that generate a line that the never touches a single coordinate. But this defeats the purpose of machine learning. You're suppose to sample data you have no knowledge about and derive a model from it.
If you can already pre-select images that form a model that never touches a coordinate, it means you already have an understanding of the model and can likely just manually tune all the weights by hand to get what you want. As regular humans, we don't have the super-human ability to do this. We can only do it for the simplest 2D example and all the higher dimensional stuff can only be understood by analogy and random sampling.
One thing you can do is just add more data to the training set. That will move the trendline and could shift it away from something it touches. But at the same time it could move it towards touching a new coordinate.
Your intuition is correct, fitting for the data will lead to this arising. However, you misunderstand the interpretation. What is being fit is exactly an encoding of the dataset (a dual of it, like switching from edges of a cube to faces), so the "coincidence" is which specific images show up with extremely high similarity after similar prompting; likely meaning nothing related showed up in the testing, so it never got pushed around and settled nearby. For many images, that encoding was lossy and sufficiently moved that they are not matches to the training data, however for some they have ended up essentially encoded in these "lines" (as your analogy goes) and can therefore be reconstructed with not that much effort.
But it's not an isomorphism. I have a function that has a range that covers every single number in existence.
Does that mean that function is A copy of every single number in existence? No. Is that function an isomorphism of every number in existence? No mathematician would use the term "isomorphism" in such a way.
You could say that if F(3) = 4, then (4, F(3)) are two things that are isomorphic. But, again, no mathematician would call F(X) isomorphic to 4.
Technically in terms of law, what this would mean is that USING the model to generate a COPY would be illegal. The model itself does not hold a copy until function application is executed. Therefore existence of the model itself IS not illegal.
To say it's illegal is to say that a copy of the mona lisa exists in paint and paint brushes. F(X) is the paint and paint brush. F(3) is the painting.
I will say this, I don't think most people in the jury will be technical enough to get this distinction and the technical part isn't even really important in my opinion. The laws should be made based off of what's better for society/humanity not what's technically valid.
Still my comment here is on what is technically happening which SHOULD remain completely separate from the politics surrounding it.
And, not to belabour the point, but a function F that has an isomorphism to the real numbers means that F does indeed encode all numbers (and, when defined as a set, exactly contains a copy of each). That's literally the point of isomorphism. To demonstrate simply, the integers have an isomorphism to the natural numbers: you count them in an alternating pattern. That's an isomorphism between them (and it need not be relation preserving). So, with that, you can interpret all integers as natural numbers; we say they're the same up to their isomorphism. Isomorphisms also exist for the real numbers with other structures, via bijective maps, and so forth.
I'm not going to respond to the other stuff since it's nothing to do with what I said or responded to (even when it was a misinterpretation).
My point of contention is that it's not a copy.
But imagine this scenario. The model produces a picture that is identical to a picture in the real world, but that picture was NOT part of the training data. This is certainly possible.
It's fuzzy but there comes a point where the layers of abstraction between two entities that are isomorphic become big enough such that they don't violate copyright law.
If a model were to produce a pre-existing image not in its dataset, then we can strongly suspect that two behaviours have occurred (for the sake of establishing provenance).
First, it was just the ever so small, non-zero chance that it was produced by a random process that was not based on anything learnt, and within the finite resolution and some accuracy of colour, happens to match up to something already existing which, for example, may have been produced independently by a human artist - say, after the dataset was collected, but before the model was trained, such that it is emphatically a "coincidence" and nothing else, no collusion, no learning, no intelligence. Just chance.
Second, and I believe far more likely, this image has been encoded into the model by indirect learning and proxy. For example, say some notable and famous work of art is not part of your dataset (which is believable). However, say that there exist, say, art that references this famous work in some way (as artists sometimes do), which may be up to the point of parody (or some a gag), or may be some small aspect (colour scheme, notable style, outfits, lighting & composition, certain people, etc). Especially if it is what we would literally call "influential", is it not possible to reconstruct the famous work (that was not in the dataset) by indirectly learning about it from other pieces? Now, exact (or near-exact) matches are unlikely, for the most part it would probably tend more towards similarly repeating the references and parody of its dataset, but I feel this would strongly increase the likelihood of randomly producing it as in the first behaviour, and we're inflating the chances certain works are reconstructed entirely by "coincidence"; at least, as according to all observers that merely look at the binary presence of the original in the dataset.
Is this a copy? Is it not a copy? In all honesty, the notion of copy or not is insufficiently defined, because there's always some angle that can be part of it, or may not be, as most copyright law is handled on a case by case basis. Is it a copy because its provenance was an attempt to copy? Was it a copy by cosmic accident or just an independent work? Was it a copy by inferred reconstruction, thus not literal copying from the source, even if a verbatim reproduction is made? Is a vector version of a raster image a copy, and vice versa? What is "sufficiently" transformative to no longer be a copy? Does a copy of the Mona Lisa exist in the mind of a person that experiences it, while as a subject in a photograph it has grounds to not be? I'm trying to avoid anthropomorphising it and state whichever way what this is, I'd rather just illustrate how what we learn, what we experience, what we produce replicas of, and all of this is generally part of an unsolved problem/field: epistemology. Do we understand diffusion models fully? No. Therefore, weighing on one side or another should merely be for the purpose of a leading research question, or to exist on either side of a dialectic; such as in the courtroom. Perhaps we're still just too immature on this topic to really say anything about copying.
Really, the only thing we do know about copying, is that you should never get caught doing it. (For legal reasons, that was a joke.)
"Copyright laundering" will be the phrase of the era. Throw Picasso in with a thousand reproductions, wash it with DeviantArt, get Picasso back out. Does it matter than it's algorithmically derived rather than a stroke-by-stroke reproduction? There's probably already ample case law around reproductions to deal with this.
y = f(x) = x + 2
This is what's memorized. But with that equation you can plug in a specific x. to get: 3 = f(1) = 1 + 2
The training set here would be (1,3). What is memorized is y = f(x) = x + 2. You can literally see that (1,3) is NOT in the model EVEN though that model CAN produce a (1,3) given the right input.I think the technical part of this is sound. People will just take the face value explanation which is I see a 3! therefore a THREE was memorized. But technically a 3 WAS not memorized. This is categorically true.
You are Right. The persuasive part of this argument is not very good though as you can see by this thread.
y = 2x + 3
And I have (x,y) = (1,5)
I'm saying the y = 2x + 3 is the ONLY thing stored in memory. The (1,5) is not stored anywhere.(1,5) is training data. y=2x+3 is the model. This makes sense.
Paint and paint brushes can be thought of as functions. You apply the right inputs (brush strokes) and you get an output that is a painting.
Therefore could you say that all paintings in the world are encoded into the paintbrush? No. You can't. Not until you use the paintbrush to copy something.
https://en.wikipedia.org/wiki/Pierre_Menard,_Author_of_the_Q...
I mean technically, if there are many copies of the same data set pair, THEN you can call it over-fitting. But removing that does not fully remove the problem. The curve will still intersect datapoints EVEN without over fitting.
Also over-fitting is not memorization.
Much of that space consists of pairs of data that aren't relevant. So to understand it in terms of the analogy of the line in 2D space... The only thing that's relevant are points near the line.
For example if we use the line to represent housing costs(Y) over time(X). You can take a bunch of housing prices sold over time and plot it on the graph as dots. Linear regression will form a line that best fits those dots. Not EVERY single section on the plane matters though. Only the dots matter and the space that's very near the trend line have anything meaningful in terms of data.
My fear is that when things like this come up for lawsuits, overconfident experts are going to talk out of their asses about how these models do or don't work, and that's going to determine how automation affects our society.
On a technical level, I'd love to see a patch-wise version of this investigation. This shows whole images being regurgitated near-exactly rarely. I expect that small part-of-the-image patches are regurgitated even more often. But is it simple stuff like edges being regurgitated or are larger parts regurgitated frequently too? Given the architectures generally used, I'd guess that it's significant.
Short answer: for any country participating in WIPO, yes it is.
Edit: And if you don't believe me, let's replace the word "painting" with the word "song":
> If you memorize a famous (recent) song and recreate that as closely as you can, is it copyright violation or a transformative derivative work?
Hopefully everyone here knows the answer is "absolutely yes!" which is why artists need permission to cover other people's music.
Now, yes, intent factors in, insofar as it affects a potential fair use defense. But that doesn't affect that status of the work, it only determines whether the action to violate copyright is defensible
This is how Google gets away with returning images in their search result, and why I can't just copy an image they return and use it without myself violating copyright.
Honestly, training a generative model on the patent database could be very useful for inventing and invalidating patents. I wouldn't be surprised if examiners are using it in 5+ years. Since they've got required response count and time, I wouldn't be surprised, if they started using it just to speed up their own work.
I wouldn't be surprised if Google is wanting the lawsuit to lose. It would block open-source models like these from existing and give them potentially a competitive advantage to be able to afford whatever compliance is mandated. They'd be able to offer services that comply, but open-source models would only have access to lower quality data and would be stunted.
'By releasing public demos that, as impressive & useful as they may be, have major flaws, established companies have less to gain & more to lose than cash-hungry startups.
If Google & Meta haven't released chatGPT-like things, it's not because they can't. It's because they won't.' (https://twitter.com/ylecun/status/1617908306420600833)
On memorization - I suspect this is a great thing for downstream performance, and a positive indicator that diffusion models are actually better generative models than prior methods (VAEs, GANs, etc). This mirrors the finding that feedforward neural networks can memorize randomly labeled data very well. Intuitively it feels like memorization is a quantifiable behavior that is a foundational activity in information processing - it is one type of optimal usage of observed data - that superpowers downstream performance.
Diffusion models are VAEs and follow the same variational framework. You could imagine that VAEs are diffusion models with a single step in the forward and backward processes ;). They actually optimize the same VLB objective, but with diffusion models the objective is a trajectory instead of a single step, however, when training we are optimizing single step transitions. This is possible because the objective ends up being a sum of logarithms, thus there is no dependence between terms.
In practice we solve a simplified objective which looks a lot like as we do with standard AutoEncoders ;)
The key component that differentiates the two is in what we expect of the underlying neural network. It is far easier to parameterize small changes than large ones, with VAEs you ask the decoder to produce a large change in the latent variable, whereas with diffusion, we generally split them into 4000 smaller changes assuming you are using the DDPM approach and not the DDIM one.
Because we are improving with very small steps, we avoid the blurriness of VAEs, and we don't go out of distribution when sampling random noise. VAEs are often difficult to synthesize because even with KLD in the objective, the encoder produces a low variance distribution, and so when we sample noise from a high(er) variance gaussian, we are out of distribution rather quickly.
I think this because (reasons I'm sure you're familiar with)
- Diffusion model is closest to a hierarchical VAE, but hierarchical VAEs were significantly less popular than regular VAEs
- The variational objective in diffusion models in practice is weighted
- Diffusion models require unchanging latent dimension while VAEs aren't restricted to this
- Historically, diffusion models grew out of score-based approaches, not from VAEs
This allows us to separate non probabilistic diffusion models e.g. cold diffusion. But then again, what's the difference between a deterministic model and sampling from a delta function? ;)
From there [3,4] show improvements to DDPMs, [5] shows that diffusion models can be very general. [7,8] show diffusion models from the view of score matching.
[1] AutoEncoding Variational Bayes
[2] Denoising Diffusion Probabilistic Models
[3] Denoising Implicit Models
[4] Improved Denoising Diffusion Probabilistic Models
[5] Cold Diffusion
[6] Deep Unsupervised Learning using Nonequilibrium Thermodynamics
[7] Generative Modeling by Estimating Gradients of the Data Distribution
[8] Score-Based Generative Modeling through Stochastic Differential Equations
[9] Variational Diffusion Models
Edit: I also want to add that google seems to try really hard to show themselves as the good guys by not releasing its models because it's not safe enough, but in this paper they used an incredible amount of computation and show me otherwise.
Also, "almost no matches" is still problematic, because they still occur, and since they don't seem prevalent, they're unlikely to have been accounted for in any other way (SD 2 being finally based on a deduplicated dataset). So, in a sense, they're "still in there", it's just harder to find by a casual or simple approach as above. Again, one of the central themes is that diffusion seems to be less resistant to attacks than, say, GANs, and that may mean more efficient and complex extraction attacks may translate over even more effectively.
Some links?
>most of their 175 million images comes from effectively "retrying" each prompt 500 times
Prompts from the most duplicated samples in the dataset, a really important aspect if you actually want to used this method in the wild, this is also one of the reasons why I said that this attack seems so implausible.
>they're usually much more targetted than this
Even if you target some images you would still need an absurd amount of luck if with the most duplicated sample you only get 109, we can be generous and think that with the whole dataset will have something like 200 matches, the probability of finding an image with a direct attack is still less than a million (even if you know the prompt) and we're not talking about a model trained on a deduped dataset.
This paper seems to answer the question of, "can SD, even just in theory, produce copyright-infringing work?" with "yes, it can."
For other images that are a product of thousands - if not millions - of source images, it becomes murkier.
Extracting images in the wild yes, the authors of the paper have access to the dataset, they could sort prompts and images based on their presence in the dataset and they have an incredible amount of computation to do so, generating 175 mln using a diffusion model is an extremely resource-intensive task.
Anyway, I don't think the point of this was to indicate that people can stumble on these incidents, but rather that it is possible. It's hard to see how this won't affect the ongoing suit.
I will move the replies here as well.
But please see https://news.ycombinator.com/item?id=34614429. This sort of surgery is time-consuming!)
I think the paper is well worth the read, it's not particularly long (much is references and appendices), and nicely written, with at least a quick bit on most things I would think to test as part of something like this. Good stuff.
(Normally I wouldn't have moved your comment as part of the merge but (a) you said something about the paper, and (b) after name-dropping metacircularity how could I not)
This was entirely predictable, and is one prong of the primary arguments that these ML models, trained on datasets including copyrighted images taken without permission, infringe on the copyright of those images' creators.
Train the damn things on public domain images and images you have explicit permission for, and you'll be fine. Stop acting like you have a right to just vacuum up every image ever created because it's "AI".
I don't agree that this is the resulting implication. I'd argue this is only implied if you believe that the software is so similar to humans that making judgements about the software is equivalent to making judgements about humans.
Put another way, when a human does those things it is not infringement, because we have already determined that it's not, and the laws are based on humans interacting with the content. A computer doing those things is arguably a new thing entirely, and requires new rules. This does not imply that the old rules would or should stop applying to humans the way they do now.
Stable Diffusion has no agency, cannot think, and cannot create on its own. A human brain, even if it has been given no art to learn about, can still create. The art influences the form of its creations, true, but in ways that are fundamentally different than ML art generators, if only because of the presence of a conscious human will actively directing it. (And no, a human writing a prompt is not the same thing either.)
People can doodle Micky on napkins or in notebooks or into their Van Gogh poster all they want. In many ways ML is just making that easier. The problems are all what you do with it, given it being so much easier, not with the capability.
Regardless, people are already, right now, using these ML art generators to cut real human artists out of the loop while producing products for commercial sale.
But however copyright law works, if little Timmy draws a Pokemon it's pretty normal and nobody gets fined or goes to jail.
The systemic effects from this are going to be because you can get good enough "original" output without needing an actual artist, not because the good enough output will inherently violate copyright.
The only reason it is able to produce art in their style reliably is because it has been trained on their work without their (or, in fact, anyone's) permission.
Computers are tools and neural networks only resemble human brains on a surface level. The human learning process is a lot more involved than just interconnections and weights, there is an entire array of biological concepts that neural networks don't even try to simulate in their approach.
The comparison between neural networks and the human brain is the same as the comparison between a hard drive and the human brain: the recollection may be automatic and similar, but the concepts behind it and the legal implications aren't.
You can generate cartoon characters without training on pictures of Mickey Mouse. Just use pictures that don't carry any copyright requirements. The tech and its many possibilities won't change. If the code is any good, the generated model will be just as good as the current one.
If a human cannot publicly use a copyrighted image without a license, why / how a non-human can?
If some image are free to use with attribution, how can an ML model track and provide such attribution?
If and when the non-human is granted human rights, this can be revisited.
Easy - if they cannot provide attribution, they cannot use the image to train an ML model.
That's a problem for the people who create the models to solve.
This is what's so frustrating about the ML/AI community, they think the onus is on everybody else to overcome problems created by their products.
And UK law is adding that right even for commercial models.
Somehow, I don't think that one's likely to fly.
https://www.linklaters.com/en/insights/blogs/digilinks/2022/...
(Note, it's looking less likely the UK one will actually get extended that way now.)
This is a great example of coming to a conclusion about copyright based on how you think the system should work vs how it actually works.
Google was able to convince the court their actions constituted fair use.
My guess is training a generative AI will also be fair use. The question is, what about the output of the resulting model? And that is a question it'll take a court to answer.
I suspect that you are right, and that it will place legal responsibility on the individual operating the software for any infringement.
I think this will force the companies building these models to behave as if training the model was also infringement, because I cannot imagine a scenario where the average end-user has enough understanding/awareness of the implications of their prompts to avoid generating infringing work, and end-users getting sued would create an instant chilling effect on the use of such software.
My bet is the court will determine that whether the output of a model is or isn't subject to copyright isn't a black-and-white answer, but rather depends on the work.
Fundamentally, the test as to whether a work represents a copyright violation is about "substantial similarity" (https://en.wikipedia.org/wiki/Substantial_similarity):
> To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright, the defendant actually copied the work, and the level of copying amounts to misappropriation. Under the doctrine of substantial similarity, a work can be found to infringe copyright even if the wording of text has been changed or visual or audible elements are altered
Okay, so let's say I take a thousand copyright images, averaged their pixels, and produced a single uniform grey output. No jury is going to conclude that work has "substantial similarity" with any of the original works, and I'm clear.
But now suppose I do the same, but weight it so 99% of the pixel colour comes from one image, and the remaining 1% comes from the rest.
Well, in that case, odds are very good a jury would find me guilty of violating the copyright of that original work that represents 99% of the image.
So my bet is the courts will conclude that the model, itself, doesn't in any way violate copyright, nor did the training itself run afoul of the law, but that any given output might, depending on the substantial similarity test.
And that means every single work is suspect and a potential target for litigation.
My understanding is that it's not difficult to tune existing models as an end-user, but to start from zero would be impossible for most individuals financially and technically.
This is how those specialized models can still generate just about anything. Without the core model mixed in, the specialized model would be nearly useless.
Anthropomorphizing this software seems problematic.
Memorising a picture using a fleshy model is fine, so the raw fact that art has been found in a black box model here isn't necessarily relevant. Might be. Might not be.
The biological/evolutionary limits of humans are core assumptions of current laws, and I’d argue that the operating environment has changed enough to make those assumptions outdated.
> Memorising a picture using a fleshy model is fine
I imagine the fleshy model is fine not because it’s fleshy, but because it’s the model that people were targeting when writing current laws.
Even if no memorization occurred, there are still big questions about why such a model should be treated like anything other than just another computer program from a legal perspective.
Likewise, in this work they prime the pump by using exact training prompts of highly duplicated training images. And then you have to generate 500 images from that prompt to find 10 duplications. You've really gotta want to find the duplicates, which indicates that these are going to be extremely rare in practice, and even more rare once the training data is hardened against the attack by deduplication.
Good analogy! An MP3 decoder takes an input and produces an output. If the output is copyrighted material, it's well understood that the inputs is simply a transformed version of that same copyrighted material and is similarly copyrighted.
The SD model is very much analogous. The prompt causes the algorithm to extract some output from the input model. If the output is copyrighted material then similarly the input model must carry a transformed version of that same copyrighted material and is therefore also subject to copyright.
Right?
By the way, I pose this, but I highly doubt this is actually how the courts will rule. I think they'll find the model itself is fine, that the training is subject to a fair use defense, but that the outputs may be subject to copyright if there's substantial similarly to an existing work in the training set.
I imagine with the right prompt one could coax out a copywritten image even if it hadn’t ever seen it before
Github Copilot could be leaking private code as well.
https://twitter.com/alexjc/status/1620466058565132288
Specifically, this post confirms cases of a single image in the training set being "memorized"
https://twitter.com/Eric_Wallace_/status/1620475626611421186...
People are generally not capable of creating realistic copies of anything straight from memory. Realistic paintings from the the Renaissance could only be done on still life, but photo realistic paintings were only possible after color photography was widespread and artists could use full scale photographs to base their paintings off of.
The average person can't even draw a bicycle from memory[0]. Gp is overestimating human capacity for reproduction.
However, whatever humans are capable of or not is a red herring. If humans memorize and reproduce a copyrighted work - it counts as a performance or copy. Someone typing Dan Browns novel on their blog from photographic memory does not get a pass, and noether should the AI models.
0. If you have 10 seconds, sketch a bicycle before opening the link https://twistedsifter.com/2016/04/artist-asks-people-to-draw...
A better example would be: could the average person accurately imagine a bicycle? Can their brain generate the correct image, even if they lack the skills to draw it?
That said, I don't think it's an excuse for these AI tools
If there is now evidence that it does memorize and emit training data, that argument had new life.
That said, I imagine it's easy enough to add a post-process step to ensure your result isn't a member of the training set.
That said, with humans it's not like anyone tries to prevent us from having the capability, it's just that usually people know better. Otherwise going to art galleries would be problematic.
Which is not the same as memorizing and reproducing a work, which is equivalent to, say, memorizing and covering/performing a song, something which is well understood to be a violation of copyright law.
> Or tattoo artists giving people tattoos of movies characters and scenes.
Which violates copyright.
> Or the multiple times in school that an art assignment involved more or less copying or reinterpretation of an existing work.
Also violates copyright, but is probably covered by a fair use defense due to the scholarship aspect of the work.
My wife works in design. It's amazingly easy to try to come up with a new logo design that somehow nearly exactly matches other existing logos that are in use by other companies. They have to spend a huge amount of time making sure their 'original work' doesn't violate someone else's copyright/trademark.
Is 'creation' the act of violating copyright? I wouldn't think so.
We tend to talk about copyright violation in light of distribution. Is the file transfer from the SD server to your browser an act of distribution? Who knows what the law thinks at this point.
> Is 'creation' the act of violating copyright? I wouldn't think so.
It absolutely can be. There's numerous cases of courts finding cases of copyright violation due to accidental copying, particularly in the music industry.
I know, that might sound unintuitive, but the law is the law:
https://www.law.uci.edu/faculty/full-time/reese/reese_innoce...
> But since 1931, a defendant’s mental state has clearly not been relevant under U.S. copyright law to the question of liability for direct copyright infringement. As the Supreme Court stated that year, “[i]ntention to infringe is not essential under the Act.” So innocent infringers are just as liable as those who infringe knowingly or recklessly.
As an aside, logo design, which you mentioned, actually comes with a whole other set of considerations, as you're less concerned about copyright and more about trademark, which is a whole different branch of IP law.
When will we learn to stop being overconfident about how these things work? Just say “we don’t know yet.” Anthropomorphism and overconfidence are dangerous in that we could set the wrong precedents (culturally and legally) for how these are used and how automation affects society.
A while back, people built a model for medical imaging that learned to distinguish between images of patients with vs without some disease (can't remember that detail). It did well, but failed in the real world. It turned out that instead of learning to recognize features of the disease at hand, it learned to recognize some tiny feature of whether the image came from a specific hospital that collected part of the dataset, or something stupid like that.
Saying "maybe model X does the same thing as humans" is proven wrong for X after X after X. At this point, the default assumption should be that ML techniques are different from humans unless proven otherwise.
You do realize that there are people in this thread who can explain to you in fine grain detail how an ML model actually comes to conclusions, without speculating abstract "weird shortcuts".
The calculation is theoretically unimportant. Practically, it is of great importance.
>>> Brains are radically different from GPUs.
>> The same calculations can be performed by an abacus. What is doing the calculation is irrelevant. The question is what the calculation is
> Does it matter if the simulations of photons is on an abacus or using a GPU? I think that's the question. Neither of those are "reality", just a simulation.
I think so, yes, specifically with respect to the question of "how do we know the human brain doesn't use the same shortcut?". Simulations likely use very different shortcuts because they're optimizing for the structural design of a man-made machine that exists today and uses numerical and CS tricks to cheapen the computation cost while maintaining error rates on training data. The brain uses physical shortcuts to minimize energy expenditure, for survival of the host, and resiliency of the species (i.e. OK if flaws exist sometimes as long as the species survival is improved long-term). So not only is ML a fun-house mirror image of a brain (our model is extremely imperfect today), the optimization process is totally alien to how the brain figured out all its shortcuts.
The irony here being we're a few layers deep in a thread started as a critique on this kind of pointless anthropomorphism.
It’s not anthropomorphizing it’s a description and an analogy.
Neural networks work similar to a brain, and it’s easier to describe them that way because, again, they were modelled that way.
It’s not a perfect analogy, but your offence would indicate you have a lack of understanding in communication, neural nets, or you’re trying to blow things out of proportion for some reason.
Surely there's a more appropriate way to say that, and a more charitable reading of my comment. As a researcher in the field, I think it's safe to say I understand the models, and maybe I am overly sensitive at people jumping to wrong conclusions because I'm so tired of it.
The issue isn't the communication aspect of the analogy, it's the reasoning aspect. For example, people who understand these things say "these work like people" (a useful analogy) and then people who don't understand them say "well if they work like people then they should be legislated like people" (not useful reasoning because the assumption in the "if" was just an analogy). The game of telephone is the danger.
You can see lay people in this very thread taking the analogies literally and extrapolating based on literal interpretations of model-brain analogies.
Eg, Someone will naively state that something is or isn’t copyright infringement “because the tool learns like humans do”… which again, is not a question that a court would ask about the tool and since copyright is a legal invention it is kind of pointless to drift off into philosophical oblivion…
There's all kinds of brains in this world from all kinds of life. A dog can learn. A cat can learn. And they don't learn like a human would.
Yes, but they learn a LOT MORE like a human would vs these machine models. Cats & dogs share the same underlying structure, from the neuron/synapse/neurotransmitter system, up to the brainstem/cerebellum/midbrain/cerebrum architecture, as well as being inextricably integrated into a living body and sensory system, and growth pattern.
And, as you say, there are big differences in how we all learn. But those differences are utterly trivial compared to the differences between humans and ML.
This probably means there are far more matches to be found that would be considered clearly copies to humans. SSIM might be a bit heavy for the task but a simple comparison of the gradients from neighboring pixels might match quite a lot more.
Signaling that you allow indexing is no opt-in for any use or abuse.
Also intrtesting: a prompt asking for a well-known image, without a watermark.
Seen from another angle, if you take a random value for each pixel, you're unlikely to generate anything resembling a picture, let alone any given training picture you're trying to reproduce. That there's an input that makes the model output a training sample doesn't necessarily mean that it's easy to find. But the paper shows that you can find several by guessing randomly.
To me, the implication is that these models cannot be seamlessly exchanged for a human brain when considering their impact and compatibility with current laws.
Of course, like an human, an NN can be used to generate copyright infringing images, but it doesn't follow that any generated image is infringing.
Most humans cannot, and this is important, because the fact that most humans cannot arguably played a role in the formulation of all current rules.
So can cameras. And there are laws that restrict where cameras can be pointed in places that do not restrict human sight, because the two kinds of "seeing" are distinctly different.
Layering AI into the mix doesn't suddenly remove or mitigate the technical realities of the software.
> but it doesn't follow that any generated image is infringing.
I agree, and this is not my claim.
Furthermore the laws tend to be a complete wreck of logical paradoxes that fall apart in situations like this, hence forcing more case law to be generated.
The reason you cannot record what's going on in a bathroom has as much to do with the implications of the recording as it does with the implications of being observed in the first place.
I don't disagree about the ensuing mess of laws, but I don't think we have a reason to believe Stable Diffusion will be spared from it.
Are you saying humans are incapable of accidentally reproducing previously seen work?
I'm arguing that a finding like this harms such core arguments, and highlights just one of many ways these models are entirely unlike humans.
> Are you saying humans are incapable of accidentally reproducing previously seen work?
To conclude that would be a form of propositional and possibly equivocation fallacy, IMO.
While I acknowledge that it's possible someone might "accidentally remember" someone else's work and then create a piece that is very similar to it, this hardly seems likely as a general case, and such a possibility was baked in to the current rules as written. To take this further and claim that a human could do so with precision and in a reproducible manner seems questionable.
The Stable Diffusion equivalent of this kind of "remembering" and resulting duplication is again different in context and contents than a human exposed to the same images.
The fact that it can be replicated systematically is the most distinctly non-human part, and is a strong hint that we're comparing very different things.
Why? If a really good artist studies a single piece long enough, he could be able reproduce it to a degree where it takes expert analysis to determine which is the original. It's not as if there have never been forgeries of expensive artworks.
The difference between a human studying a certain piece intensely, and a model overfitting to it, is that to the model, it happens by accident. Overfitting to the training set is not a desired outcome, it's something ML techniques are trying to actively avoid.
Part of the answer is in your question. A really good artist is rare, and forgery is a specialized skill.
The point is not that humans are capable of forgery, but that the model/software is capable of it with minimal effort, and enables anyone regardless of skill to achieve similar outcomes.
Setting aside the resulting privacy issues, this has major implications. Two are:
1. The risk of forgery previously accepted by a copyright framework that assumed human actors has drastically changed.
2. The claims that Stable Diffusion only produces derivative work and not originals was at the heart of the argument that says SD model training is fair use, and regardless of why, we know this is not true.
> The difference between a human studying a certain piece intensely, and a model overfitting to it, is that to the model, it happens by accident.
That is one difference in the factors surrounding the initial creation of the output.
The differences continue to stack up as you examine the process of learning, the algorithms that produce new output, the computational context (both hardware and software), the agency of the entity creating the output, etc.
The software doesn’t do anything by accident. It does exactly what it has been instructed to do.
The humans training the model are at the core of any accident. This might seem obvious, but I think it bears restating because it highlights the nature of the situation. The model is not an independent actor.
Whether or not it’s a desired outcome is not relevant. The fact that it is possible - and that this possibility has further implications about the nature of the software itself - is what I argue hurts the standard arguments claiming Stable Diffusion is not infringing by virtue of its similarity to human processes of thinking and expression.
People who could machine very small cogwheels precisely enough used to be rare, which is, as far as I know, part of the reason Charles Babbage never completed the Analytical Engine (prohibitive costs for the parts). Today we mass produce tiny cogwheels and other mechanical parts due to automation.
> The claims that Stable Diffusion only produces derivative work and not originals was at the heart of the argument that says SD model training is fair use, and regardless of why, we know this is not true.
Brushes, canvas and pigments can be used to produce non-derivative works as well. So can pencils, or photocopiers. There are thousands of tools that could be used to infringe copyrights or make forgeries of other peoples works. We don't blame the car when bankrobbers use it as a getaway vehicle.
I don't think an analogy that compares the manufacture of machinery with the creation of artwork is a good one. I understand the similarity you're teasing at, but I think this is a category error.
> Brushes, canvas and pigments can be used to produce non-derivative works as well.
None of these tools require the original artwork of other artists to function. They are primitive tools, and I don't think it's reasonable to include them in the same category as an AI that was explicitly trained for a particular purpose.
> We don't blame the car when bankrobbers use it as a getaway vehicle.
We might blame the car manufacturer though if the car was autonomous and had been trained on a dataset of all drivers and driving styles, and had "learned" to be a getaway car, replete with situational awareness of a typical robbery scene, knowledge of how to avoid cops, etc.
I am generally concerned, these days, because it seems like the days when I could just keep my head down and do science is over, and now I have to also defend myself against opportunist hype-masters who probably were jumping on the crypto bandwagon six months ago and now are self proclaimed AI experts.
Will you expand on why? This is a big accusation to just throw out without exposition, your disagreement with the notion of a "lossy db" notwithstanding.
You seem to be projecting quite a lot with this comment, re: snark.
A few moments searching quickly show Alex isn't some crypto bandwagoner, and the implication that he is is distinctly unhelpful when trying to judge the value of each participant's comments in this conversation, here and on Twitter.
It is now clear that SD is treading on thin ice: training on watermarked and copyrighted images without their author's permission, then attempting to commercialize it even when the model emits images that resemble a high similarity of the original training data including watermarks or copyrighted images: (Mickey Mouse, Getty Images watermarks, Bloodborne art cover, etc).
This weakens their fair use argument, especially with Getty Images also threatening to sue SD for the same reason. If OpenAI was able to get permission to train on shutterstock images [1], then SD could have done the same, but chose not to.
Perhaps SD thought they could get away with it and launch their grift (DreamStudio) on digital images and artists. It turns out that now SD creates an opt-out system afterwards but artists can already find out if their images are in the training set. [2].
[0] https://arxiv.org/pdf/2212.03860.pdf
[1] https://www.prnewswire.com/news-releases/shutterstock-partne...
I wonder whether extraction attacks easier if you have many ancestral models?
[2] https://github.com/AUTOMATIC1111/stable-diffusion-webui#stab...
It's a surprisingly common experience for music students to excitedly tell everyone they know about a new piece of music they've been composing, saying it is probably the best thing they've ever written, and then a friend or teacher has to say, "I don't know how to break it to you, but you've 'composed' the Xth movement of Beethoven's Yth symphony."
And sometimes they will say, "I have? I don't think I've ever heard Beethoven's Yth symphony." But of course they have, just without realizing it. It was in the background of some movie they watched or something like that.
Unlike humans, I don't think AIs have any belief about whether their work is original or not, but it's the same type of error. And with similar legal consequences: people have been sued for stealing a melody (presumably not always consciously). The difference with AIs is they can produce much more output than humans, and it's muddier what is actually doing the creating (AI authors? users?).
not really, no. It's more the case that a NN is an encoding of a large dataset space, and if you go far enough away from the average you may find yourself on an island with an individual outlier. This is why you simply get back the training example, it's regurgitating all that it was fed.
For example, try to train a model using a single image, with a space as wide as a sole point you will find that all you get back is that point.
A human being re-inventing without remembering the source is distinct from remembering the source and returning it when lacking other data.
Of course then a company in China (or other countries that are rich enough but don't care the U.S. lawsuits enough) will make a better SD 3.0.
If the neural network tech itself is good enough, this shouldn't be too much of a problem. Finding these images and verifying their copyright status is going to take more work.
Most likely, AI giants will buy digital art websites like DeviantArt or Flickr and use the images already collected after changing around the terms a bunch of times.
To be honest, I already feel quite strange about folks profiting off of SD results.
A lot of that existing content is copyrighted, and demonstrably present in the resulting product; so these companies have an uphill battle _not_ to be stopped by the courts.
These lawsuits will go exactly like in the war against piracy, impossible to win. Cut one head off, ten more appear.
I'm not arguing that generative art will cease to exist, I'm just arguing that the path to legal revenue will involve paying copyright holders; and the real gold in this industry is in licensing training datasets to businesses that just want 'AI Magic' in their product marketing.
It's good news for the artists and programmers whose work was being copied, though.
Let me put it another way. Suppose you authored and copyrighted two images. I then train a model that can interpolate between those images using a "latent space" from 0-1.0. For most of that space, the image might look totally different, or even like noise. But if I set the value to 1.0, generate your exact image (pixel-for-pixel) and then sell it, would you say "well since model could have generated different images, the fact that it generated my image isn't infringement...". This is about like saying that taking a digital photograph and then selling the photograph isn't infringement because the pixels are slightly different, and the digital camera is merely a model that approximates the true colors of the atoms by sampling photons which it maps into a regular grid to approximate colors, which are further down-sampled by means of compression algorithms. Truly, the digital photo and real-life image can't be considered the same, can they? The digital approach only reproduces it using "heuristics" and "algorithms", and could theoretically reproduce any image given a slightly different set of input bytes to the decompression, rendering, and printing algorithms.
After all that's exactly what my computer does in case of images that it downloaded from the internet which I open up later from local drive to view them again.
Take another example: if you try to sell or publicly play a Taylor Swift song without permission, that is illegal (even though it was technically generated by your computer and you used an arbitrary process to create the reproduction), its owned by Taylor Swift. Its long been the case that the song is reproduced by dozens of different mediums, storage procedures, algorithms, sound generators, and mechanisms, but ultimately it is still the creative property of Swift.
But more seriously: I think the question is perhaps a bit vague.
Of course google image search can show a wrong source, but it's still a big difference between "absolutely nothing" and "potentially wrong thing".
From where I stand, it's a cheap-suit excuse to bleed us of culture and art.
Of course maybe in the future we'll have a better method, but it will be a very different thing from what we have today.
Compression is prediction.
"Critics claim that models such as Stable Diffusion act like modern collage tools, recreating copyrighted and sensitive material.
Yet, our new paper shows that this behaviour is exceedingly rare, recreating copies in less than 0,00006% of 175M test cases."
* being unable to recover many or most images does not imply that you will be unable to recover any at all. Being able to recover a nonzero proportion is still a problem even if the proportion is small.
* the existence of artifacts in the recovered images is not in itself sufficient to prevent legal claims, and yet ignoring them is plenty sufficient for increasing the proportion of inputs that can be considered recovered
When in reality, you absolutely cannot store data to reproduce every image it's trained on because the model isn't big enough to hold even a fraction of what it's trained on.
You make a claim that seems trivially false. Cite some example comments to back it up.
I mean, it literally says "some" right there in the title, and if you click through and actually read the paper you see it's a tiny proportion.
Here is a real world example.
To make a really dumb example:
I can make a black box model that I push in 1000 bestseller novels through, that has the total "storage size" of 1 novel. The pigeonhole principle says my model can't possibly contain all 1000 novels. If you ask it (anything) it will respond with that one novel, verbatim. It never reproduces anything else.
Does it now matter whether I trained it on 1000 novels? Does it matter that my "black box" is just that, a literal black cardboard box containing just a copy of a novel?