AI is in danger of being swallowed up by copyright law
heathermeeker.com
heathermeeker.com
The fact of the matter is that the AI companies don't want to ask for permission, because people will say no. Or worse, ask for attribution or even payment. There is plenty of copyright free/public domain material out there, but what the customers of AI people want isn't available under those terms.
The code to train an AI is not enough to make a product and these people have nothing to add themselves, so they take what others made and use that to make a profit. They can make or pay for their own paintings, their own pictures, their own music, but that would require putting in too much work or paying too much money.
It's very possible that a judge will rule that AI models do not violate copyright. If that is the case, I hope new legislation will correct that oversight very quickly.
You know that Stable Diffusion lawsuit? Go check who the lawyers behind that work for; Disney wants that same outcome.
If a company invests $250 million into an original movie, I don't see why they shouldn't have some say over their content for at least a couple of years. Not until 2150 or whatever the end date for modern works is supposed to be, but give it some time at least.
OpenAI is the result of billions being thrown around. When it comes to billions, it doesn't matter if they come from Disney, Google, Microsoft or Amazon. None of these companies have our individual rights at heart, they only care about profits.
In this rare occasion, the interests of the people and Disney align. The laws protecting the independent writers/programmers/artists are the same ones that protect Disney.
The tools themselves work on arbitrary data sets. Anyone who can dig up enough public domain/attribution free pictures/code/text can train their own AI without even coming close to copyright issues. Hell, had these super smart AI people managed to find out a method of attribution, the data set would include massive amounts of works released under Creative Commons or open source licenses.
Good quality data and more of it means better output. Disney is almost certainly doing their own thing internally, benefiting from their ability to use both free as well as their own IP and the capital to hire cheap workers to train it directly.
It's not that I don't understand why artists might be upset about a company scraping copyrighted art, I just think that the longer term effects of legally kneecapping open source variants while handing over the most powerful versions of it to the existing intellectual property giants are A Bad Thing.
It’s also worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar to some other existing artwork (ie, how all art works through reference, training, and inspiration). A world where Getty and Disney control even more than they already do.
2. People using non-copyleft license just do it because public domain seems to have a complicated legal status across the world.
Then, a lot of open source authors understand that their works are not groundbraking inventions and want to share them with the world without any fees or costs. And others have groundbraking inventions and still share them for free with the world.
But there are almost always license terms attached to the piece of the work. They can be essentially non-limiting, like public domain or MIT. But they can also enforce some minimal requirements like attribution. Why should any entity, especially huge corporations, not be bound to those conditions? It was mostly those corporations which created a restrictive copyright. Try to draw an image of Micky Mouse and put it on your website and see what happens.
The simple fact of the matter is that Disney and Getty invest a lot of money into these materials being out there in the first place. Open source programmers and artists spend a lot of time producing works for no cost other than some minor courtesies.
AI companies aren't your friend or the little mom ''n pop shop down the road. Their technology giants backed by billionaires. When it comes to Disney versus Google/Microsoft, I'm against both sides if it means giving up my rights.
Big AI taking your stuff and ignoring copyright law isn't some kind of protests against copyright, it's the very opposite; it shows that copyright doesn't matter if you have the money to defend yourself in court. Violate the the MPAA's copyright and you get extradited, violate some random person's copyright and you should feel honoured that people even want to steal your work.
In my opinion, the idea behind the current copyright system works fine if the terms weren't so ridiculously long. Restrict copyright to five or ten years and I'd be fine with the whole thing. This "70 years after the death of the author" crap is the biggest stifle on copyright adds.
It's Disney that are, by proxy, suing Stable Diffusion to create the legal precedent that you desire.
I firmly believe in "practice what you preach". I you declare you firmly believe in A but then do something directly counter to that because it's more convenient in this specific case, then that doesn't sit right with me.
Besides, further expanding copyright in this one area will only make it so much harder to reduce it later. And the pro-copyright folks will be able to say "you say you want less copyright, but you vigorously advocated in favour of copyright then, you hypocrite!" (and they wouldn't be entirely wrong, either). All this effort and energy fighting ML tools would be better directed at reducing copyright instead.
I don't disagree with your view on corporations. Do I like what CoPilot is doing? Not really. But at the end of the day: does CoPilot's or ChatGPT's mere existence really take away anything concrete from me? Am I harmed or even inconvenienced by it? Are my rights reduced? Is my code harmed by it? Is my income reduced? I don't really see how it concretely affects me, other than a general "feeling of unfairness".
And I see real risks with all of this: most regular people and small businesses don't have the resources to litigate as it's expensive and time-consuming, so a "license" that you or I slap on a piece of code is, realistically speaking, just ink on a piece of paper. GPL violations are rampant, violations of other licenses probably happen even more (but people generally care less about that, so not as widely publicized). Who will benefit with more copyright law on their side? The ones with deep pockets and many lawyers on retainer. i.e., the corporations neither of us like. Think creative new copyright lawsuits such "we claim copyright on the Java API" kind of stuff.
They aren't training it on Microsoft or GitHub code.
> A world where Getty and Disney control even more than they already do.
This is exactly what is currently happening though, it's okay to rip off the little guy artist or coder. The argument here is that one big guy stood on another big guys foot, and as the little folks we shouldn't stand for it either.
I can’t speak for everyone, but personally I find that copyright can be used properly or abused, at both sides (holder/consumer). It doesn’t mean that copyright is bad, only particular caregories of claims and usage are. But abusing copyrighted material from millions of little creators at insanely automated scale is another level of evil, especially when they explicitly require consent for exactly this type of use.
worrying that requiring consent to train an AI model will inevitably lead to requiring consent to make handmade art that’s a little too similar to some other existing artwork
That’s the root of misunderstanding, afaict. We can agree that at-scale processing is bad and that fair use is still okay. A human with a pen (or a text editor) can’t damage copyright at scale by learning terabytes of material in few weeks and producing the same amount in hours, so they can be excluded from this. Humans who use AI can, so they’re a target.
The copyright terms means that for Life+70 countries only works where a) the author died before 1953 and where the works were published before 1928 are in the public domain.
An AI training on a given work should comply with the law and with copyrights, just like anyone else. It should also respect the license or other terms the works were released under. -- You could easily silo the data by license, and have a different model per license.
Patents should be a good thing (they allowed inventions to be published instead of being kept secret). However, it is easy for large companies to get patents on trivial things, write overly broad patents, and collate a large number of patents in a domain. That means trying to innovate or compete in a highly patented field like audio or video compression is difficult.
Or you can (correctly) think it's a huge drag on innovation and human progress.
If you think the latter then hoping in this case for legal precedent to broaden the scope of copyright enforcement is just bizarre logic. This isn't a rule that already exists as such. The case will set a precedent (based on interpretation of existing law) for the future.
When a new technology is introduced, for example the compact disc was invented, lawyers get to poke whether that "distribute" applies to the music CDs, or just to vinyls and music tapes (because at the time of granting that license, CDs weren't yet a thing! gotcha!).
The answer to this conundrum might vary in different countries, and we can have fun discussing that in the context of AI, but it does not affect how handmade art shouldn't be too similar.
Do you ask for permission when you train your mind on copyrighted books? Or observe paintings? Or listen to music? Do you ask for permission when you get new ideas from HN that aren't your own?
Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity.
Time to put the horse back in the barn, cars and trains are here.
Fundamentally the current accommodation of copyright has two main justifications: 1) Protect economic activity 2) Moral right to identify original author of a work
The point of ease of replication speaks to (1) fundamentally breaking. A human can only produce so much output compared to an AI system.
(2) is a much thornier subject, and not one I really feel qualified to speak on.
Yes, and humans are being found liable for copyright infringement for doing so. All that's needed to establish liability is access and substantial similarity; the bar for the latter can be very low indeed (see Williams et al. v. Bridgeport Music et al.).
I’m not able to read billions of books in less than an hour.
Even if we agree that machine learning is like human learning, scale commonly matters in law.
I think you underestimate the sheer volume of data + conclusions the brain ingests and processes on a daily basis, primarily through unconscious experience.
While our brains are more complex than the networks, this has never been in dispute.
The quantity of experiences needed to train GPT-3, however, is many more than we are capable of experiencing in a lifetime.
Also, about volume of data processed by the brain: https://gwern.net/Differences
But I think you meant GPT-3 has seen many books during training, not during inference. You should know that training on millions of books is not the only way GPT-3 learns. It is just the foundation of its knowledge.
GPT-3 learns "in-context", that means it can learn a new word or a new task at first sight. It just needs a description or a few examples. This is the most powerful feature of GPT-3 - in-context learning. And when it comes to ICL, it is much like humans - only sees a few examples, not millions of books.
> “Do you ask for permission when you train your mind on copyrighted books?”
The nature of ICL is that it happens at prediction time. So GPT-3 would have to explicitly be instructed to learn a specific skill. Should it reject instructions if they are sourced from copyrighted books?
I’m not a lawyer, but to me it seems within the realm of possibility that a U.S. court eventually finds strongly in favor of the copyright holders, the Supreme Court agrees (because Big Tech has so few friends left), and OpenAI will be required to destroy the GPT-3 model and all copies of the training data because they can’t filter out copyrighted works.
Just because you can find one way that GPT might be slow doesn't invalidate the point that its training does use massive amounts of data.
I am not a lawyer, so the following is only my opinion.
Scale matters, but so does the legality of the thing that scales.
Reading two dozen books by other authors, or studying hundreds of artworks, or visiting the museum of awesome statues every week, in order to get inspired for ones own novel/painting/scuplture, isn't illegal.
So a lawsuit will have a really hard time argueing that it somehow is a problem if its two dozen billion books/paintings/sculptures. Because, such a lawsuit would suddely need to explain why the smaller scale is also problematic, only less so. And given that this is basically how art worked ever since the first human had the idea to paint pictures on a cave wall, that's a hard sell.
Humans either learn art by being natural art geniuses, or by receiving instruction and learning through an iterative process (where, again, they might create thousands of art works, but nowhere near the scale here), which is very different.
Why is it different?
The only difference that matters, is scale. And again, if I want to argue that something done 10000000000 times is legally problematic, I have to be prepared to explain why doing it 10 times is problematic as well, only less so.
I am completely aware that scale is a "thing" in legal systems. But as I said before: For scale to be important, the unscaled act in itself has to be problematic already.
Granted, things around power concentration have deep philosophical and social roots.
Could I have a dollar? What about a billion dollars?
The burden isn't on me to explain why something being scaled up by a billion is not the same.
2. An AI has a training set of every image in the world and produces an entirely unique work.
Which do you have more of a problem with?
Copyright law serves the purpose of peoples works being protected from unauthorized parties making copies of their works, and profiting off them.
It doesn't protect from technology making the production of new works cheaper, faster, more efficient. An artist using photoshop can be, and is allowed to be, many times faster than one using oil and canvas.
I think the objectionable thing about the second is that the AI knows everyone's styles, and so can use them in creating something new. Even if the AI is restricted to not be able to paint an image in a certain artist style (as the new version of stable diffusion is, for instance) and the art is unique, I think part of the problem is that the AI is still (presumably) leaning on the collective styles of everyone it has trained over.
If we can train an AI over a small dataset or maybe even a large dataset of old art, or some mix in between, and then maybe fine tune it wtih a snall sampling of modern art, then I believe it would be unobjectionable, as this is largely how humans do it.
ML can do either facsimile or imitation far better than a human mind can.
You seem to be conflating both things and suggesting that ML only does facsimile, which is where the potential legal problems are.
No, it is not. Memorization =! understanding.
I can teach a parrot to spew the times table, good luck getting it to understand how to apply it.
And a parrot is billions upon billions of times more capable than any current AI algos.
Diffusion would appear to me to work in much the same way. It doesn't understand what's good ("works", creates acceptable output) or why, but it knows it when it sees it, and has the tools to refine it.
AI, like humans, is capable of both imitation and facsimile. It is far superior at both feats.
Your fallacy is that you are noticing AI is superior at facsimile and erroneously assuming it is “not learning”. You are also ignoring the other amazing learning feats of imitation in front of you.
Parrots are lovely animals, but it’s unclear what you think you’ve accomplished by bringing them up. The fact that they are capable of more than just memorization does not differentiate them from advanced AI models, which are rapidly gaining all sorts of abilities.
Polly the parrot would have a hard time producing a picture of Elmo with a light saber in a Superman costume riding a dragon on the moon in the style of Rembrandt (in under 300ms, at least). I also know a parrot couldn’t write a 500 word story about the image.
I am pretty sure that GPT-3 is a lot more capable than a parrot in transpiling a function written in Python to Golang, or writing a summary to a tech-magazine article.
Same as StableDiffusion is ALOT more capable than me in drawing, painting and generally making up pretty pictures.
But there are no copies. For example, the LAION-2b training data is a total of 240 TB. The pruned SD model based on this dataset, is less than 5GB.
The data isn't copied into the models, it is used to teach the models, letting them learn patterns in the dataset.
Whether or not it's identical to human brains isn't the matter, they'd need to prove how a small 5GB model trained from a huge dataset infringes their rights specifically.
I have a 1.9 GB mp4 file on my harddrive. It contains 2 hours and 15 minutes of 1080p video data at 24 fps. Assuming it was generated from 4096x2160 16-bit color depth source material, the "training data" was 10.32 TB. I bet I could even get a similar size reduction as LAION-2b if I recompressed it to 720p.
Could I not also claim that I created an advanced AI model, which did not copy but learned patterns in the dataset? Modern video compression algorithms are getting quite complicated, after all.
I think no reasonable person would agree with this, but can you prove that the AI model is doing something substantially different?
Such patterns would enable the video file to decode into a multitude of pictures not originally in the training data. Obviously, a video file cannot do that...it's just compressed data.
Generative models however can generate things that are not in its training set.
And of course, there is a fundamental difference in the source data between compressed video and a generative model: video codecs work with a sorted sequence of images, where most images are slight variations of the ones before them. The training for generative AI doesn't have these properties, the input is not an ordered sequence, and even similar pictures are not sequential variations of one another.
Relying solely "uncompressed" size does not a really good metric make (this is analogous to the raw input size of the LAION dataset): one could make a reasonable argument that there are not billions (1) of image-pairs that are effectively identical up to a minute shift. I would posit the correct basis would be the Shannon entropy of the "best fit" ordering (minimize inter-frame diff), versus the lossy-compressed video, and a similar "best fit" ordering for the LAION dataset vs. the model.
My suspicion is that one will find that the relative number of "smooth transition" pairs in LAION viz the whole will be very different from the video.
-------
(1) - Napkin math: There are about 194400 frames, so ~37 billion (37791165600) frame-pairs. Assuming you have runs of about 1 second between hard cuts throughout, so an incidence rate of 1/24 for non-smooth transitions, gives us about ~36 billion "smooth transition" frame-pairs. I think it is safe to assume "on the order of" 1 billion, then. This ignores long "action" scenes with significant variance in images throughout, but also ignores longer-than-1-second slower scenes, hence the order-of-magnitude shrink in the assumption as buffer.
No, but what if you then produce "your own" rendition, or "remix", of that book, song or movie and offer it to the public? E.g. you memorize a collection of Taylor Swift's latest songs, and then start performing a medley of her hits in your local clubs, you may well find yourself in trouble.
But that never, ever, ever happens with current AI. it is not AGI. it is inspired by nothing, has no creativity, nada, ziltch.
Or it’s just too fast so let’s stop it?
Bear in mind, AI is not making artists or creative types obsolete - that would be fair game, just like computers made human calculators obsolete. No, this is about abusing other people's work.
Where is the abuse happening?
If using someone else work for learning is infringement, then that's going to cause a lot of difficulty for all artists. Try making a rock song without listening to rock, or paint some modern art without viewing it etc.
The loophole is to use copyrighted works for free despite no learning taking place - no human being observing and developing their skills based on that work - rather, an algorithm transforming those works into some other useful interpretation of them.
[0]: https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In.... [1]: https://www.vox.com/2018/11/7/18073200/aereo
I don't know if that counts as plagiarism, but there's clearly some use of this copyright material that the authors probably didn't envision and did not grant permission for. I have no idea what the law would be in cases like this
the data was originally permitted to be copied.
The question isn't whether the training is violating copyright - as long as the data set had permission to be viewed (which it must have, since it was public).
The question is whether the final result - the model/weights - is a derivative work of the training data set. If it is a derivative work, then the model must be in violation of copyright. But copyright law allows for sufficiently transformative work to be considered new, rather than derivative. So is training a model using methods like this constitute a transformative work?
This was about the humans consuming other people's content.
> Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity.
If humans make stuff that is too close to someone else's source materials then it is considered plagiarism and not "inspired by".
> For any image that the AI generates, you can't point to any image in the training data that the image is derived from.
Why can't you point to the Getty Images watermark that it is quite happy to reproduce? Isn't that surely evidence that it doesn't actually understand what it is reproducing?
> The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work.
I have also seen billions of images, therefore I cannot be actually store the real images in my head and thus nothing I paint could ever be considered plagiarism. That's brilliant, I think there are a few law firms defending artists who would be looking to hire you.
Also, the fact that Artistic Freedom is now under attack by artists. Not that long ago Artists hated the Music Industry and Corporations such as Disney for weaponizing Copyright law against Artistic Freedom. Now artists are utilizing that same tactic against other artists.
What is your point? If "Linux users" are right it is not FUD or the concept of "FUD" is pointless.
you don't need permission to train on books, but you do need to buy the books or take them from the library one at a time.
"training" these machines so far is not like human learning as becomes apparent when they spit out source code that mirrors individual repositories. And you know that humans are required to both remix their own creations and follow copyright law at thbowe same time, and also adhere to the social and institutional stigmas against extensive uncreative cut and paste paraphrasals.
when training AIs on copyright law trains them in obeying copyright law, they'll be ready for the Turing test, or even to be called AIs.
That's not a problem, we already have copyright laws that prevent people from distributing mirrors of copyrighted works. They don't care about how the works were copied.
yes, isn't that what very many people are saying?
Humans are capable of both facsimile and imitation.
The fact that ML is able to perform facsimile far better than a human can is not evidence that this is “not the same” learning. Only that ML learning is superior. ML is far superior in feats of both imitation and facsimile.
If I show a 3 year old a single picture of a Tiger, and tell him this is a tiger, the child is able to recognize a Tiger fairly accurate in real life without further input. Though the child might say that a house cat is tiger,,,,
ML learning needs millions of pictures to do the same, and still might mistake an elephant for a tiger...
ML is nothing more than graph approximation, there is no logical reasoning
ML is currently capable of the tiger case you mention. It’s generally called “few-shot” or “one-shot” learning. In the context of an image generation model, having never seen a tiger before, if you show it a few pictures of a tiger, it could immediately draw you thousands of tigers in any variation or scenario you can think of, which is way more than a child can do.
As for the need to train on millions of images for the base model, I believe you are trying to say something about “sample efficiency”, and how ML differs from the brain in this regard outside of the few/one-shot contexts (which ML is absolutely capable of). I would argue that sample efficiency of the brain is actually also quite low, much lower than people assume. It’s irrelevant to an argument that ML is not superior, because ML is clearly is capable of learning richer, more effective representations in a shorter wall time than we can, whether it is sample efficient or not. And in the sample efficient few/one-shot contexts (learning what a tiger looks like from one picture), it also outperforms humans in speed accuracy and creativity. It’s not even close.
As for classification errors, ML is capable of some errors we are not, actually by virtue of being superior at learning representations we are not even close to being capable of learning. But those are edge cases, and they are fixed by various means. In the main cases, ML outperforms humans in speed, accuracy and class complexity, all exponentially.
You said something about graph approximation but it doesn’t make a lot of sense. I’m talking about learning and you’re complaining that machine learning is not “logical reasoning”. Whether ML is currently capable of logical reasoning is another discussion. Certain models do demonstrate some types of it today.
“Graph approximation” is a type of learning task. ML is a billion times better than humans at it so it also doesn’t help you argue that ML isn’t superior (in that regard).
Yes, you do need to buy books, which gives you permission to read them.
The thing that authors are trying to argue here is that they should get to control what type of entity should be allowed to view the work they purchased. It's the same as going "you bought my book, but now that I know you're a communist, I think the courts should ban you from reading it".
No, that's not it. It's more like if I memorized a bunch of pop-songs, then performed a composition of my own whose second verse was a straight lift of a song by Madonna. I would owe her performance royalties. And I would be obliged to reproduce her copyright notice, so that my audience would know that if they pull the same stunt, they're on the hook for royalties too.
Now, moving from holding the model creator culpable to the user would obviously be problematic as well, since they have no way of knowing whether the output is novel or a copy paste. Some sort of filter would seem to be the solution, it should disregard output that exactly or almost exactly matches any input.
Github ignored the licenses of countless repos and simply took everything posted publicly for training. They didn't care whether it was available to them entirely legally, they just pretended that copyright doesn't exist for them.
Other licenses such as the MIT license require that you name the original creator.
A license allows new uses that copyright would otherwise block. Some kinds of AI training are fully local and don't make the AI into a derivative work, so they don't need any attribution and you don't need to accept the license to distribute.
It's not obvious to me that the implicit permission we've been granting for humans to view our content for free also means that we've given permission for AI models to be trained on that data. You don't automatically have the right to take my content and do whatever you like with it.
I have a small inconsequential blog. I intended to make that material available for people to read for free, but I did not have (but should have had!) the foresight to think that companies would take my content, store it somewhere else, and use it for training their models.
At some point I'll be putting up an explicit message on my blog denying permission to use for ML training purposes, unless the model being trained is some appropriately open-sourced and available model that benefits everyone.
actually you don't have the right to restrict the content, except as part of what's allowed in copyright law (those rights a spelt out - like distribution, broadcasting publicly, making derivative works).
specifically, you cannot have the right to restrict me from reading the works, and learning from it.
Imagine a hypothetical scenario - i bought your book, and counted the words and letters to compile some sort of index/table, and published that. Not a very interesting work, but it is transformative, and thus, you do not own copyright to my index/table. You cannot even prevent me from doing the counting and publishing.
The section titled "Exclusive rights in copyrighted works".
There are 6 rights.
(1) to reproduce the copyrighted work in copies or phonorecords;
(2) to prepare derivative works based upon the copyrighted work;
(3) to distribute copies or phonorecords of the copyrighted work to the public by sale or other transfer of ownership, or by rental, lease, or lending;
(4) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and motion pictures and other audiovisual works, to perform the copyrighted work publicly;
(5) in the case of literary, musical, dramatic, and choreographic works, pantomimes, and pictorial, graphic, or sculptural works, including the individual images of a motion picture or other audiovisual work, to display the copyrighted work publicly; and
(6) in the case of sound recordings, to perform the copyrighted work publicly by means of a digital audio transmission.
It is not illegal to read a stolen book, only to steal the book.
I pay the books directly (cash, credit) or indirectly (school books via taxes). I do pay the louvre to observe the painting. I also pay to listen music in ads (YouTube) or via subscription (YT Music and Spotify).
This whole thread really makes me want to pull my hair out.
Difference between illegaly creating a (even temporary) copy of a copyrighted work (e.g. streaming a movie) vs. creating a derivative work of said copyrighted work: Two completely different things, with completely different legal outcomes.
If OpenAI in any shape or form creates a temporary copy (<--- by copyright definition of what a copy is!) than this needs to be adressed with the former. If OpenAI creates a work that is considered to be a derivative work (<---- by copyright definition of what a derivative work is!) than that needs to be adressed with the latter.
The crux of this whole thing is: Human minds cannot make a copy of a copyrighted work by definition of copyright laws (in Germany, I presume the same can be said for pretty much all western copyright laws), while anything that a computer does can be construed as making a copy.
but that's not the point of contention. The training data set has been granted the right to be distributed (by virtue of it being available for viewing already - it's not hidden or secret). The proof is that a human can already view it manually. Let's call this 'public'.
The question is, whether using this public training dataset constitutes creating a derivative work. Is the ML model sufficiently transformative, that the ML model is itself a new work and thus does not fall under the copyright of the original dataset?
This is wrong. My paintings are publicly available (especially going by your definition [which I'm confused by the origin of?]). Taking a photograph of my paintings is still a copyright violation. I hope we can ignore all the legal kerfuffle about personal use, as it has no bearing on our discussion. Again -- all of this boils down back to what I've said before: Bare human consumption does not constitute as making a copy, nearly everything else does.
Your second point -- a copyrighted work automatically granting someone else any rights (especially distrubtionial rights) by just being available to be consumed -- is even more wrong. I'm not going to go further into that, as you can very easily prove yourself wrong by googling it.
>The question is, whether using this public training dataset constitutes creating a derivative work
I'm not well versed in the US copyright laws, but I would assume (strongly so) that this would not be the case. I -- again, for US copyright law -- assume that for something to be considered a derivative work, it needs to include (or be present in other ways) copyrightable (!) parts of the original work(s). In other words, the original work needs to "shine through" the derivative work, in one way or the other. The delta of parameter changes of a ML model would (imo) not constitute such a thing.
Problems with derivative works will come into play when considering the things ML models produce.
You are mixing up the two things that I've mentioned in my original comment. You have to differentiate between creating a copy and creating a derivative work. Both of those things matter, when talking about AI, but the former is way more cut clear.
>The question is - to what extent does the exact image of your painting remain within the AI's data matrices?
And the answer is: It's irrelevant. The model has to be ingested with a copy of something. That's all that matters. The AI could even reject learning from that something. By the time that something reaches the AI to even do something with it, it's been copied (in the literal sense) who knows how many times, each of those times being a copyright violation.
I would put the same criteria to the copy made for the purpose of AI training. As long as you have the right to view the image, you would also have the right to ingest that image using an algorithm.
Most people looking at AI can tell it is not like a mind or an artist, because of certain intuitive arguments which boil down to their surprising ability and their bizarre faults (drawing hands is still a struggle for most models) and current limits (you have to hack prompts instead of asking naturally due to their current limits). You can reason about people using these arguments because they are people, but you cannot use it when applying it to NN because they are not.
I'd argue that the moment you start using the "people are AIs" argument, and you are then implying the converse is true "AIs are people," then you are assuming that there is some bidirectional here, and thus other qualities you assign to people, like, "people have rights" and "people deserve to be paid for their labor" and "people have rights to the work of their own hands" then must apply to AIs. And, therefore, the AI tools you are using deserve to be treated with the respect and dignity you had to treat artists and developers with before, and thus, should be paid for the work they create. That is, if they learned and created art in the same ways people create art. Just as we do not have that nursery does not own the art a child born there creates, or a university doesn't own the art an artist who studied there creates, you cannot make the argument that the work an AI creates is that of the "owner" or "trainer" of a model unless you are arguing slavery is in fact okay in this day and age. All of this of course hinges on the supposition that AIs are people, and that they learn as people do.
So, you cannot have it both ways. You cannot keep treating AIs as people in your arguments, but then deny them agency that is due to people. The only way this is is that you deep down do not believe they are people, or that you think people do not deserve rights or deserve compensation for the work of their own hands.
Also animals can be trained and make outputs and nobody accuses them of copyright infringement. That's a much better analogy here than leaping to the idea of treating one of these models like a human.
Ok, I just have to link this here; https://youtu.be/dKFunwOzEos?t=711
The difference is that I buy books, pay for visiting museums and buy music in several formats, or pay it accepting to receive advertisement between songs.
It Is expected that If I buy a book I will be allowed to read it without asking for a permission.
What I don't do is copypasting paragraphs of other books to write a new book and claim that is mine. Is a different situation.
There is zero creativity, zero art, zero original thought, zero newness.
When actual AGI happens, then your arguments mean something. Such as in, at least 50 or 100 years down the road.
Try to make a car using the best ideas developed by each car maker. You will be surprised.
Copyright law is not about logically perfect system, but creating a general environment in which artistic, academic and other creations can appear and benefit the general population.
Yes ... and because it's not a logically perfect system, its lifetime has to be limited. One day we should abolish copyright and find a better, more functional way to drive progress.
Where goes wrong, is when individuals and cooperation believe that such monopolies should be indefinite, and pushed the monopolies beyond the lifetime of the author. A dead author can’t produce new works, so it’s now clear how allowing such long monopolies increases the amount of creative work produced.
The original primary objective of copyright was to create an environment to could produce an endless supply of public work, freely available to all. It’s only abuses of copyright over the past 50 years that have destroyed objective, and ironically it’s copyright holders like Disney that really starting to suffer the consequences.
Winding back copyright durations to better balance the public and private interests would go a long way to resolving many of our issues with copyright today.
> Yes ... and because it's not a logically perfect system, its lifetime has to be limited.
It’s also worth pointing out that no system of law is “perfectly logical”. It’s almost certainly impossible to produce a perfectly logic system because humans are inherently illogical, and binding them into a perfectly logical system of law would almost certainly produce more injustices.
Anything that has infinite supply and zero marginal costs, as Nobel Prize winning economist Samuelson argues when he was looking at the context through lighthouses[0], should be free to all. By using copyright to make it a monopoly and allowing the extraction of monopoly rents you are drastically reducing the value and reach of the thing that was discovered. Copyright is a hack and this hack is now fundamentally breaking. Instead of trying to save the hack, we need a full rewrite. If winding back the duration of copyright is correct, the best winding back is zero.
As we are a remix culture where idea A and idea B combine to create idea C, we drastically reduce the innovation in our economy through reduced discoveries. This failure ends up with large monopoly holders consolidating into bigger and bigger entities in order to right some of this failure, but that only makes the monopoly extraction worse.
The discoverer should be subsidized for the discovery of that information but it should immediately go to the public domain. How you work out what that works out to is just as abstract as what Spotify works out what each play costs. This is no doubt monstrously complex to figure out the dollar number what some discovery is worth, but it is the economically correct path. Copyright isn't.
[0]: https://courses.cit.cornell.edu/econ335/out/lighthouse.pdf - page 359, first paragraph
No, obviously not, as that would clearly be unworkable and ridiculous. Mainly because we have a very unclear understanding of human creativity, and there’s no way to analyse an individuals mind to understand how they created an idea. Additionally copyrights reach generally stops at the point of “transformation”, once you take any idea an transform it “enough” it’s considered a new idea.
The reason none of the above applies to AI is simply because we’ve declared that only humans can transform and produce new ideas. AI aren’t human, thus they’re not afforded the same rights. Arguing about if there’s an inherent difference between AI creations and human creations is pointless, the law doesn’t care, it has already declared that there’s a difference between AI and human.
If you disagree with declaration, then you need to lobby to change the law. But until the change occurs, your believes are meaningless in the eyes of the law.
In my opinion things like selecting a training dataset and then writing prompts are not creative processes they are mechanical processes. Input in, output out, with barely any interaction from the human.
Consider when you commission an artist to make a painting. You give them a "prompt" by explaining what you want. Maybe you even give a "training dataset", a few examples similar to the look and feel of the result you want.
Then they go off and make something. They show you in process stuff and you make suggestions so the next version they show you is closer to what you want. This repeats until you are both happy. Then you own the drawing. Because you paid them for it.
In this case however it's absolutely clear that you did not create the work. You had input into the creation, but the artist was not a tool you are using to realize your own creative vision. They are the creator, you're a customer for them.
I think this is a specious analogy at best. The two are remarkably different contexts. AI can work at a significantly greater rate. There's also a very large question about whether for profit commercial software should be afforded the same leeway we give to ordinary human behaviour.
If an AI reproduce a copyrighted work they should then be sent to robot jail, and the human who requested the work should be sentenced for conspiracy to commit copyright infringement.
It might however be a bit early to let the horse back in the barn.
Why go down the route of turbocharging new forms of rent seeking?
Plenty of people have been successfully sued if their work is too similar to existing content.
This isn’t a new concept that AI is throwing into contention, it’s literally just companies trying to side step copyright law because of “disruption”.
Source: I work for a company in this field and we do gain permission from creators before training our models on their content. It’s very possible to operate this way but a lot of companies simply choose not to.
Little do some of them know that OpenAI was able to get permission from Shutterstock via a partnership to use their copyrighted images in the training set for DALL-E 2. [0] There is also a reason why Dance Diffusion was trained on only public domain music and copyrighted music which has the actual permission from the authors. [1] If they did otherwise and monetized on copyrighted music without the permission from musicians or record labels, they would be sued to the ground.
With the recent cases of Getty, Shutterstock, and even as admitted by the CEO of Stability themselves [2], the way forward for using copyrighted images in the training set for commercial purposes, is via licensing. Neither Getty or Shutterstock are looking for banning it, despite the AI bros claiming that these companies are trying to.
If not, just train only on public domain images to avoid these legal issues.
[0] https://www.shutterstock.com/press/20435
[1] https://techcrunch.com/2022/10/07/ai-music-generator-dance-d...
[2] https://twitter.com/EMostaque/status/1603390169192833027
Where in the guidelines does it mention that one cannot say 'tech bro, finance bro, pharma bro, and more recently and most actively the crypto bro'? These have been there for years despite the guidelines existing.
Me saying 'AI bros' is no different. Given it is fine to mention the tech bros, finance bros, crypto bros and the other, then it is also fine to say 'AI bros'.
yes, thats why I pay a fee to buy/borrow one (or someone pays the fee in the case of a library.)
> listen to music
again money is exchanged.
> Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity.
yes, and so long as they are not derived works, its not a problem.
Copyright is there to allow you and me to develop things and make money from it. It is there to stop people stealing our work, which may have taken years to develop and sell it for a profit with none of the risk.
Large corperations have abused this to make monster profits.
Google have spent billions to try and persuade us that copyright is evil, because they didn't want to pay content producers to host their work (ie music and movies on youtube and local news site)
The issue is this, I might have made a website that tells users how to make a specific type of metal work. I have a free ebook, and I run courses. I have spent many years to to perfect the art, create the tutoring content, recording videos. its advertising supported, and people are asked to consider buying a course, to support the creator.
The AI company comes along and scrapes all the content, allows people to regurgitate it, with more or less accuracy.
The creator now gets less traffic, less money and now cant afford to create more content.
The AI people now skim all the money, and the consumer gets less useful information.
Culture isn't free. Someone is paying for it, and if you stop paying them, then it doesn't get created.
As parent said: Everything is derived work. We are remix machines. It is how we learn and how we make money. Now with AI, apparently, we are offended, when something does it better and faster than we can? To me it seems, if we expect AI to pay additional fees, the question is: Why?
I am not saying that it's not an important question. Google has built its entire business around information other people have provided. I would argue most people are quite happy with the existence of something like Google search and see it as a net positive in their lives. Does that make the business part okay? Where do we stand on this in regard to an open web? Is it okay for Google to do what they do (and if they do it well to win the space), or should there maybe be a license where people have to pay the owner whenever they are indexing a website? I don't know. Feels complicated.
> If you stop paying them, then it doesn't get created.
That's an interesting thought. But is it true and, more so, is it a problem? What if humans from here on will only be paid to create stuff that an AI can't?
I look forward to a life of horrific poverty
If AI takes jobs because it's simply superior at them, and that creates friction and anxiety until we have stuff figured out, that's of course sad and we should do our best to soften the process, but I think it's inevitable. The carriage must die. It seems obvious that restrictions on training data are just a distraction and will not move the needle on any interesting time frame.
If however AI does not pay forward, in an arrangement that makes our collective lifes better, I will be the first to work on burning it into to the fucking ground.
But, on a lighter note, since that has generally been the direction of human civilization (not linear when zoomed in, but always when zooming out) I remain optimistic.
The weavers were left to rot when the automatic looms came in. (there were in flanders east england and northern france incredibly rich and influential class)
Furniture makers were left to rot when steam power tools came in
Farm labourer were left to starve when steam threshing/harvesting came in.
enclosure was another tragic note in england.
The green shirts were lobbying for "a share of the domestic profit" in the 20s-30s, in the 60s they were convinced that we were going to be working 2 hours a day by now, with robot servants cooking and cleaning for us, and no-one would be living in poverty. Even Orwell has written on this.
Instead we see productivity in the western[2] world dropping. Meaning for every human hour worked we make less money. because I suspect in part to the rise of servant-as-a-service jobs(food/shopping delivery/cleaning/elderly care etc etc) all of which are long hours and low paid.
[1] well, DDR everyone had a job, but lived in permapovety and were likely to be disappeared if you spoke out.
[2] specifically the US and UK, who appear to be snorting financial inequality by the metric fuckton
What I was more so thinking of are the unspecific societal functions that evolved to the benefit of everybody, but more so to those who could not have afforded them beforehand: Quality health care, various forms of social support, more accessible education and food, better road systems. The stuff that makes the charts on education, prosperity and health go from bottom left to top right and child mortality and hunger in the opposite direction.
The injustices of the day do not show in the most important, most long term graphs. As far as I can tell (and I am happy to hear your thoughts) this can only be true because people have benefitted increasingly from things improving, over time.
If you're referring to piracy, that is very much being kept in check. Otherwise, the vast majority of copyrighted art is only available for payment in various ways (streaming services, museum and theatre access fees, library cards, buying e-books etc).
someone pays, just maybe not you. How do you think google/meta/et al offer you a service free at the point of delivery, through charity?
> You can google an image of any great work of art and look at it for as long as you like
see my bit about google. The copyright still is with the owner. That image can be removed, should the owner wish, but for various reasons its too expensive to get google to respect that.
> I would argue most people are quite happy with the existence of something like Google search
yes, because its a symbiotic relationship. I as a creator, make something that people want to find, google points them to me, and I get people's attention. I might do that to fluff my ego, or try and convert it to cash through sales or something.
The AI step threatens to remove that relationship. Instead of being passed to me, the AI just pastes shit its gleaned from mine and other websites, leaving no chance of me getting a reward for making that website.
If the copyright on a given work of art is still active, those pictures were taken and are distributed with the permission of the copyright holder (or they're just pirated). That's one of the reasons it's much easier to find images of classic art (for which the copyright has expired) than it is to find images of contemporary art.
> You don't need to pay for or "borrow" anything to learn from copyrighted works.
What exactly do you think copyright is, and would you be surprised to learn that libraries have purchased the books on their shelves?
That doesn't seem to be universally true, but an end-game of capitalism. There are countless examples of artistry/sculpture/music that were created long before copyright existed and although they may have been "paid" for it previously, those cultural items can be appreciated without needing to pay someone for it.
There are also many contemporary cultural items that were created without monetary recompense that can also be enjoyed without needing to spend money.
> Copyright is there to allow you and me to develop things and make money from it. It is there to stop people stealing our work, which may have taken years to develop and sell it for a profit with none of the risk.
Your use of the word "stealing" is unnecessarily loaded and specifically means that the creator was deprived of physical ownership which would be incorrect.
You are arguing against your own point here. As I said culture stops being created when there is no money for people to create it.
Should copyright never expire? no. is 25 years enough? you betcha.
> There are also many contemporary cultural items that were created without monetary recompense
Again you are missing the wider point. For culture to be created you need a mix of people, and those people to feel safe enough, and have enough time and energy to create said culture. They will also need money for materials.
As I suspect you are not on a poverty wage, you will have the time, energy and healthcare to be able to create a new thing. This is not a luxury someone who works two jobs just to make rent has.
> our use of the word "stealing" is unnecessarily loaded and specifically means that the creator was deprived of physical ownership which would be incorrect.
stealing is taking with intent to deprive. I mean specifically what I say.
taking someone else's work and selling it as your own to make money, whilst depriving that person of credit or income stream. It is morally wrong.
Now there is an argument about corporations abusing copyright (they do) but, throwing it all out only benefits people like google, amazon and facebook.
This is just deeply wrong. Culture existed before money. It is tragic to me that a person can't see culture as anything but a marketable good.
with respect, thats not what I am saying, I'm saying it has a cost. If people do not have the means to spend that money on making culture, then it is not created.
Juvenal was a client of someone, and complained about it
Tallis, Allegri, Purcell, Bach, Mozart were all professional composers
The great seats of learning (Ashurbanipal's libary, Venice, Alexandria) are all paid for by a ruler wanting to show off how good they were
Wilde, byron were all rich people wafting around bored and making art on the way.
In the 60s-80s it was possible to live in NYC working at a bar or something, and still have time and money to create art. Where can you do that now?
Now you need to be rich, or have time, or get patrons. The internet is a great way to either lower the cost of entry (see music) or get support to create (see Patreon)
> This is just deeply wrong. Culture existed before money
Culture existed when we had time, food and resources to stop worrying about being cold wet and hungry.
Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted.
> Do you ask for permission when you get new ideas from HN that aren't your own?
You don’t need to. It’s implicitly assumed, by virtue of publishing in a public forum, that the author is providing permission for people read their comments and ideas, and remix them as they wish. That permission doesn’t include exact replication, but reading and understanding is assumed, otherwise why did the author publish it?
> Humans are constantly ingesting gobs of "copyrighted" insights that they eventually remix into their own creations without necessarily reimbursing the original source(s) of their creativity.
Correct. Literally everything produced by a human is automatically copyrighted. But the manner in which work is published creates implicit licenses for the public to consume those works. You publish in public, you automatically grate licenses for the public to consume and transform it.
If a human transforms an idea, it automatically becomes a new idea with its own copyright. The same doesn’t apply to AI because they’re not human, and thus the law generally doesn’t recognise them an having ability to create or transform ideas. If you believe AI can create and transform ideas, then you need lobby for the law to recognise that ability, but right now, only natural humans have that ability according to the law
No you don’t. That would fall under the category of “derivative work” which is still the intellectual property of the original author under most jurisdiction copyright laws.
Therefore, using a training dataset does not constitute copyright violation.
If the AI outputted an exact copy (or a close enough copy, that the laymen would agree it's a copy), then that particular instance of the AI's output is in violation of copyright. The AI model itself violate any copyright.
> Therefore, using a training dataset does not constitute copyright violation.
It's not for you to decide that. Different jurisdictions will have their own process for deciding that and none of them are based on the opinions of random commentators on internet message boards.
Also please bare in mind my comment was reply to a specific statement (repeated below) and not talking about AI in general:
> You publish in public, you automatically grate licenses for the public to consume and transform it.
^ this statement is not correct for the reasons I posted. AI discussions might add colour to the debate but it doesn't alter the incorrectness of the above statement.
> If the AI outputted an exact copy (or a close enough copy, that the laymen would agree it's a copy), then that particular instance of the AI's output is in violation of copyright. The AI model itself violate any copyright.
That assumption needs testing in courts.
As I've posted elsewhere, there have been plenty of cases where copyright holders have successfully sued other creators based on new works that have bared a resemblance to existing works. It happens all the time. I remember reading a story about how a newly successful author was being handed ideas from fans during a book signing only for one of her representatives to intercept them each time. When they later asked why the representative took them, the representative said "it's because if any of your future books follow a similar idea, that fan could sue. But if we can prove you haven't read the idea then the fan has no claim". (to paraphrase)
Experts don't all agree on where the line is with similar works created by humans, let alone the implications of copyrighted content being used as training data for computers. And this is true for every jurisdiction I've researched. So to have random people on HN talk as confidently as they do about this being all perfectly legal is rather preposterous. You don't even fully grasp the intricacies of copyright law in your own jurisdiction, let alone the wider world. In fact this is such a blurred line that I wouldn't be surprised if the some cases would have different rulings in different courts within that same jurisdiction. It's definitely not as clear cut as you allude to.
If we agree on this, what we need to resolve mostly seems to be, in how far a human should not be allowed to use publicly available data to make his tool, in the same way he is allowed to use publicly available data to make anything else.
When you buy a book, you’re not paying a licensing fee. You’re exchanging for goods. You’re granted very few rights to own a copy of the work. But they’re almost all to do with distribution. None of those rights is the right to read it.
>You publish in public, you automatically grate licenses for the public to consume and transform it.
By this interpretation, all the artists upset by stable diffusion have given tacit permission for their works to be used as they are published in the public. Even though those works are posted to websites, the artist has not granted any rights to the viewer of the work.
> only natural humans have that ability according to the law
The law is not explicit about this, and we have case law that describes non-human entities as having rights associated historically with personhood. This is definitely not clear, nor is it obvious.
It's not just assumed, it's celebrated when a work of art gathers fans who produce their own, inspired content.
Not sure why it needs to be over-complicated or different for silicone neural nets. But I think it will get very over-complicated, if not politicised, in the following years.
It is implied that if you are using the work by youself, or via a tool you made youself, it's fine.
However works that you redistribute, by copying it yourself or indirectly by tools, said silicone neural nets being one example, instead require a "wide redistribution license agreement", and those are implicitly limited by default unless the work is put in a sorta public domain license.
You are absolutely buying a license to read the material when you purchase a book. That's why books cost more than the paper they're printed on and why pirated books are illegal. The "distribution" rights you refer to stem from the "first sale" doctrine[0], which acknowledges that the first sale (e.g., you buying a new copy of a book) of a physical object embodying a copyrighted work grants limited distribution rights.
Following this logic, isn't training AI on Github or Deviantart 100% fair game then? It's not like OpenAI is infiltrating computers and reading hidden away data.
Personally I reject that. ML needs to be restricted heavily.
Sounds like an option instead of a need.
Forklifts are "agents of humans" but you still need a license to drive one.
It's pretty obvious to me at least that AI bros are using these tools recklessly and inappropriately, without regard for licensing or copyright, and therefore I am proposing that the tools need to be regulated.
Simple as that.
Unlike forum comments, GitHub code generally has an explicit license attached which you'd have to respect - you know, for instance by giving attribution to every MIT-licensed source that was used.
And even then, let's say someone releases a book with all your HN comments: you are definitely entitled to sue them for copyright. Here's some info from the BBS era, which is still relevant today: https://www.templetons.com/brad/copymyths.html
a similar ruling will also be a disaster for software as our tools of expression are very restricted. code is based on boolean algebra and predicate calculus, practice guides like design patterns and books teaching algorithms and data structures.
there are lots of ways to write bad code and only a few for good, correct code. Recognizing this led me to replicating known working code, code I had created, for multiple employers. so who's copyright did I intentionally violate?
I think we are attacking the wrong problem WRT ML and copyright. to me, ML shows the foundation on which copyright is built is a lie. we should use ML to break copyright for code.
For example, the "clean-room design" method of copying a work exists precisely to avoid potential copyright issues. One team reads the original work and writes a description in such a way that it cannot possibly be infringing, and a second team reads the description and creates the new work. This avoids any chance of someone reading the original work and incorporating potentially infringing aspects into the new work.
> Yes, that’s exactly what happens when you buy a book, or pay for a music subscription. The work is in the public domain, then global permission to observe and copy the work is already granted.
You can buy a book, read it, sell the book, and then write and sell another book based on the ideas contained in the first book (Baker v Seldon). This is the cornerstone of contemporary copyright law. Or read the book on a shelf of a bookstore where the clerk is asleep. Or borrow the book from the library or any other manner where direct compensation of the author is nowhere to be seen.
Copyright is consistently interpreted in alignment with the needs of public learning, both by protecting the authorial incentive as well as protecting the public need for knowledge.
>>You don’t need to. It’s implicitly assumed, by virtue of publishing in a public forum, that the author is providing permission for people read their comments and ideas, and remix them as they wish.
Ideas are not eligible for copyright protection.
Libraries exist.
My experience with ML tools like co-pilot is why I reject copyright claims on ML systems. there are a tool that generated original work based on my instructions not unlike a paintbrush, photoshop, or a CNC machine. My instructions were based on my exposure to copyrighted works.
I use co-pilot as an accessibility device enabling me to write code again. like with speech recognition co-pilot is a force multiplier IF you change how you work. If you keep using the habits formed by typing, you will get shit results.
The end result of the shift in how I work is now I know how to tell co-pilot how to write code in my style. My co-pilot generated code is no less my code than what I generate by hand. Co-pilot acts as an extension of my brain, not my fingers.
Is my co-pilot generated code copyrightable? I say yes because it is the result of this human's creation and instruction.
The law already makes many distinctions between humans and machines. For example, looking out the window to see when your neighbor is going to the supermarket: allowed; using a machine-vision system to store the movements of groups of people into a large database: not allowed.
Also, "training the mind" and "training a machine learning system" are two completely different things, even though the language used is the same.
It seems to me that one side is arguing that people (as in, individual human beings) already do what the AI is being accused of, the other side argues that it's replicating work.
The truth of the matter is that what is taking place is a different thing altogether. We do generally deal in a different way with "machine behavior" because we recognize it being automatic and reproducible matters.
Of course not, and given my ability to train my mind on thousands of books in a few minutes and spit out a full book based on that training in whatever style one wants in a few minutes for that as well, it seems especially unfair that people act as though there might be a difference between the two situations.
This comment fundamentally and dangerously misunderstands Copyright Law. Insights are not copyrighted, nor are they copyrightable. Copyright law controls who gets to distribute a specific “fixation” or performance of work. It is not, and never was about preventing the spread of ideas. Authors and artists have always intended for you to read/observe/listen to their work when you legally acquire a copy. They just want you to not copy it verbatim, but go do your own original work if you want to distribute or sell something.
The whole problem is that today’s NNs are specifically designed to remember and remix only the fixed performative parts of the work, and they, unlike humans, don’t understand the insights at all. They are just deterministic machines that copy and remix at a large scale. As such, it’s pretty clear the people training AI today should expect to have to ask permission before “training” (copying) other people’s work.
If you’re executing a NN algorithm in your mind, or via pen & paper, then you are copying from the training samples, because that’s what the algorithm does. During training you compute errors against the samples, and update your weights to reduce error. During inference or generation, you use the weights (the results you remembered across all your training data) to produce an output. When your training samples are clustered in the latent space, the network will only remember an average of the samples, but samples that are sparse and don’t have close neighbors are sometimes remembered verbatim because there’s nothing nearby to average from. You can legally run the algorithm all you want on your own. Once you run it and then distribute the output, it might be in violation of Copyright Law if you accidentally reproduced one of the samples. Same is true for traditional human learning, you can free copy ideas legally, but reproducing too closely something that someone else made may be against the law, even if it was accidental.
Thoughts are never illegal wrt US Copyright Law. It’s a straw man to insist on making this point.
> In other words, the end user of the model is the one to be held responsible if they reproduce and distribute the copyrighted material.
No, this is false because it is the creators of the model that 1) did not legally acquire the source material and 2) distributed the network that contains latent copies of the source material that end users can use to reproduce works from.
This is incorrect. As another poster mentioned, it is not illegal to read a stolen book. It is only illegal to steal the book.
Secondly the source material is acquired legally since it is open to consumption on the open internet.
Thirdly model does not contain “latent copies of the source material”. By using a simple test (currently legal standard) that if I showed you the node weights and counts of the network no person even trained in the art can identify it to a specific piece of work. Therefore it is at best a derivative, reasonably distinct.
Nope, this is strawman and continuing to demonstrate a misunderstanding of Copyright Law. There is no such legal standard, where did you get that? If the network can reproduce a work, then it does in fact contain a latent copy. Arguing that you can’t see it by inspecting node weights is straw man. You cannot argue that you’re not copying music if you use a new compression algorithm and then suggest it’s distinct and derivative because nobody can read the raw compressed data. That’s not how Copyright Law works. If you can approximately re-perform someone else’s work, you’re in violation. This is true even if you have to run a black-box program to produce the output.
> no person even trained in the art can identify it to a specific piece of work
Ironically, you’re actually admitting that even AI researchers can’t prove the network won’t reproduce someone’s work.
The rest you seem to now be looking for a snarky gotcha, which if you don’t want to have a discussion, then I’m uninterested in discussing further. I made clear above and in a sibling comment that remixes are gray area, and this question is complicated. That said, even if AI people do acquire source material legally, they are in fact copying it and distributing it, and that part alone can potentially violate US Copyright Law. This isn’t even up for debate, so I don’t know why you’re attempting to suggest otherwise. The lawsuits mentioned in the article were brought on evidence that networks violated copyrights of specific existing works, and lots of people have found specific examples of violations.
1) the creating of the model is does not violate copyright. Claiming otherwise means running same algorithm in meatspace would violate copyright laws, which implies thoughts violates laws which is absurd.
2) distribution of the model does not violate copyright laws because the models themselves do not contain latent copies of the work. The model itself is not the work nor a recognizable copy of it nor can it be reconstituted back to the work. It is a tool more analogous to photoshop where the tool can be used to reproduce copyrighted work, yes, by the end user (where I believe the responsibility lies). But the tool itself is not copyrighted work. Microsoft word can be used to generate copyrighted books if I’m correct. Or I can hire smarter tool: a human writer to produce copyrighted works. Is the writer-for-hire illegal? Or his employability is illegal? Of course not. I believe the law will eventually take the position that AI model is a tool.
> nor can it be reconstituted back to the work
This is false. It has already happened multiple times that networks reproduced copyrighted material.
Secondly you seem to be conflating the “tool itself” to “what the tool can do” to be strongly equivalent. I.e if the tool has the capability to violate laws, then the existence and distribution of the tool itself also violates said law. (Not so)
> if the tool has the capability to violate laws, then the existence and distribution of the tool itself violates said law.
That’s right if you remove the word “existence”. Distribution of a NN model that violates copyright by reproducing copyrighted works is illegal. That part has been my point in this thread, it seems like you understand now and we agree.
It’s “existence” is not illegal under US Copyright Law unless you didn’t have the legal right to use the training material, and in that case it’s illegal to use the material whether you used a computer or your brain, it doesn’t matter how you created the neural network (or even whether you created a neural network), the violation there isn’t the act of creating the network, it’s the act of stealing and using material you don’t have permission to use.
This whole discussion would be a lot less frustrating for you if instead of making assumptions and logic arguments about brains and computers, you took some time to read the copyright legal code. https://www.copyright.gov/
Cars, phones, guns, knives (practically anything) can be used to generate activities that break the law. They are perfectly legal to distribute. The onus on the legality of the activity lies with the end user.
> If you understand that the model is a tool, and that as a tool it can be used to generate activity that can violate laws and be used for other perfectly legal activities, then as a broad principle the distribution of said tool is not a violation of said laws
That statement is incorrect, the logic is flawed. Just because a tool has both legal and illegal uses does not necessarily have any bearing whatsoever on whether the tool’s distribution is legal. Tools that are illegal to distribute can have legal uses, and that does not make them legal to distribute.
Making statements and assuming the truth without reason nor evidence nor examples to back it up. Logical fallacy of begging the question. You have also not reasoned how freely available information is “illegal” to read/index/store amongst other things.
Not here to win you over. The audience can see how weak your position is. My last response here.
A silly example. Making GPT write a rap battle between Keynes and Mises goes beyond a performative remix, it is transformational work, nothing is copied explicitly. If a human were to write it that would not violate copyright.
I think that to tackle this we need a new lens other than copyright in the long term.
The argument that NNs aren’t memorizing is definitely debatable and not necessarily true. They are designed to memorize deltas and averages from examples. They are, at the most fundamental level, building high dimensional splines to approximate their training data, and intentionally trying to minimize the error between the output and the examples. It’s fair to say that “usually” they don’t remember any single training sample, but it’s very easy for NNs to accidentally remember outliers verbatim. The whole reason the lawsuits mentioned in the article are happening is because we keep finding more and more examples where the network has reproduced someone’s specific work in large part. If we’re going to claim that today’s AI is producing original work, then we have to guarantee it, not just assert that it doesn’t usually happen.
> a rap battle between Keynes and Mises goes beyond a performative remix, it is a transformational work, nothing is copied explicitly.
I don’t buy that the work can be called transformational just because the remix doesn’t have any recognizable snippets. GPT is in fact copying individual words explicitly, and it’s putting words together by studying the statistical occurrence of words in context of other words.
> I think that to tackle this we need a new lens other than copyright
I totally agree with that. This question is legitimately hard. We do need a new lens, but we might have to keep and respect the old one too at the same time. I feel like AI work should acknowledge that difficulty and step up to lead the curation of training sets that are legal wrt copyright by design, rather than ignoring the concerns of the very people who made the work they are leveraging.
Training AI without permission is sneaking a camera into an art gallery without permission.
AI is not a mind. It’s a program. We might call it a “mind” as a metaphor, but it’s not really one.
So any justification which presupposes that an AI should be able to do something (really: that the people who are running the AI programs should be able do something) because they are a “mind” is fallacious and doesn’t need to be interrogated.
AI is not a mind. A mind is a physical object, a brain inside a skull inside a person. An AI is a computer program.
And while a nerd who forgot how grass feels like might confuse the two, the courts won't.
Are we really going to play devil's advocate so much that we consider these early day A"I" tools as equivalent to humans? I personally have absolutely 0 qualms about treating humans and these ML tools as completely separate entities governed by completely different laws. AI SHOULD be heavily restricted, we're already headed not towards any sort of apocalyptic singularity, but a singularity of pure, endless spam spewing forth from every orifice of the internet and elsewhere.
If these megacorps behind this AI push want it to succeed, then they should be paying for access to the images/texts/music/videos/whatever they're trying to harvest en masse. I couldn't care less if an AI learns the same way a human does or any other anthropomorphising the AI crowd want to gaslight everyoen with.
But "eating" is a fun word.
What the customers of AI want is accurate predictions of the models, and they can get that even if everyone demanding to get removed from the training set would be removed.
The makers of generative AI could remove every living artist who wants to from the dataset, the model would still develop a general solution of color theory, composition, almost every artstyle in existence, ... because fact of the matter is, there is just that much data out there. Our species collectively has spend DECADES recording, storing and categorizing everything and the proverbial kitchen sink. There are god-knows-how-many petabytes of data available in images alone, so even if just 1% of that could be used to train generative models, it would still be more than adequate.
And soon after that, there is an explosion of new generated art, filtered through the aesthetic sense of millions of humans, that can just be fed back into the models, to make them better.
The end result is the same: High-quality image generation on a scale hitherto unseen, running even on consumer grade hardware. And what lawsuits will be filed then?
I swear when I see this argument because it makes me angry.
You’re right, but they didnt, because they were too lazy and cheap to do it that way.
…and that’s why people are angry, and rightly so. Fully licensed models are the future, and it’s both irritating and disappointing that we are where we are right now because the people training these models were too lazy to assemble a training dataset that wasn’t problematic (ie. full of porn and copyrighted material).
You can argue the “but at the end of the day it’s all the same…” argument if you like, but clearly from the lawsuits it isn’t ok
They’ve completely messed it up.
There’s a reason the openai api terms of service says that “the Content may be used to improve and train models”; they’re setting themselves up to have a concrete defence for the source training data for their models.
Good job.
Stability can burn in a fire. They’ve really trashed the reputation of generative AI in a way that is going to be very difficult to recover from.
Well, a lawsuit isn't a decision, we will have to wait for the courts to decide wheter it's legally okay or not.
Reputation damage has been done.
Undoing that is going to take time and effort which, could be spent on more productive things.
I’m disappointed in where we are right now. It was entirely avoidable.
Lazy. Cheap.
/me shakes head…
That reputation damage you think matters doesn’t exist.
Somehow we have managed to come full circle to the first episode of HBO's Silicon Valley.
If a model was to add attributions to each of its answers, then perhaps the search engine analogy would hold. But, they don't (and right now, to my understanding, can't.)
A search index usually links to the source. Without that a search index is worthless, you can't use content if you don't even know where it comes from and who holds the rights.
Google search links to sources like Wikipedia in its info boxes, because without that you can't know whether the info is reliable or sourced from my brother's coworker's imaginary flat-earther friend.
Would that mean you can simply use one AI (or more) from anyone else to train another AI?
Of course access can always be limited to an API with rate limits and per-request costs, which would make it difficult to straight up copy the whole thing, but it would be hard to justify any legal protections against it.
But I want to argue here that for purposes of this latter question, your proposal of copyright enforcement (or anything similar) is too little to late.
-These "copyright violating" AIs have demonstrated the proof of concept and the damage is done. Even if these AIs are banned, the companies will just parallel reconstruct it by running the 80/20 rule: pay tiny amounts to get most of the data. After all the creators of the data were doing it for free and are in such fierce competition there's no bargaining power.
- More nefarious AIs will just do transfer learning on intermediate neurons, very difficult to prove stealing here.
- Even if you get the system to work, what about future artists and writers? Are we just creating an entrenched historical group of creatives getting royalties forever?
The distributional problem is not well solved by copyright, and better solved with e.g. corporate taxes, income taxes, VATs.
The boat has long since sailed on this… ands it’s globally entrenched as a norm of international trade that we are all “ok with this” regime of 75 years or century plus copyright terms …
And arguably the entire copyright vs AI/ML training datasets debate is founded on the notion that the artists individual copyright will last long enough that it’s going to outlive the average artist. If we look at one of the old copyright regimes, for comparison… in a world where copyright is a short default/implicit/automatic term (14 or 28 years) and the copyright owner can elect to register and pay for extensions (for a more modern twist, preferably combined with increasing incentive to prevent perpetual renewal abuses by Disney, et al)… now imagine how much data from up to 28 years ago there is, the catalogue of art and photographs and text and books and academic writings… all public domain because the authors didn’t consider them of sufficient value… all free for the ML model training… this gets even larger with a 14 year term…
Suffice to say that we are seeing systemic impacts already, culturally we’re seeing more and more money put behind less and less content controlled by fewer and fewer people due to a slow death spiral off copyright stranglehold across multiple industries, written, visual, audio and video arts are all dominated by large corporations holding IP … yes individuals continue to create, but other than rare breakthrough chance successes and internet age viral success (which are often just completely arbitrarily/random and have no real quality) these companies decide what will be popular culture…
My prediction is that the AI/ML models will be allowed but heavily scrutinised, under the simple legal doctrine that the user is the one committing the infringement since the primary purpose of these models is not infringement but unique creation, but suspicion will linger by artists and it will become a normal part of contracts in the art word…effectively an artist equivalent of the way police in many places view spray cans… just as the primary purpose of spray paint is not to create illegal graffiti, which is the justification many places used to overturn poorly justified civic bans on possession of spray paint.
I’d like to see any more draconian spread of derivative work rights (style rights etc) to be accompanied by drastic reductions in the automatic copyright term, as the ability to churn out lots of automatic content drastically lowers the value of long long terms, and the counter argument that it makes the existing rights more valuable is fucking insane as we do not need to pass copyright down to the great-great-great-great-grandchildren… the terms are already too long.
Personally I’d like to see the right to train statistical models on any works without the permission of the author enshrined in statute and an end to common-law copyright, a return to the Statute of Anne 14/28 time length, and a clear delineation between the “work” as having an author for an eternity but having a “copyright of the work” vastly limited in scope.
Ask yourself, do we want to be extending the reach of large copyright holders like Disney into taking a fee from LLM producers because they COULD be helping people draw Mickey ears on their private creations?
This is Betamax all over again and luckily that Supreme Court opinion will favor heavily in the lower court’s judgement of these models as fair use.
The flip side of this is that if we undermine paid creators until there's no incentive for them to create, then the AIs abilities stagnate on old data and we as a society drop or at least diminish the skillsets that could create new media.
AI can generate stuff humans care to look at only because of the availability of data that humans created for eachother to enjoy. As tastes, fashions, zeitgeists and pop culture change amongst humans the AI models will always be behind and unable to follow trends completely. I think.
The incentive to create is almost never financial. How many artists finance their creative efforts by working day jobs? Making a living as an artist is more about buying yourself the time to focus on making art than it is about making money. People will continue to create art, however they can, because they must.
This is all stuff I am actively thinking about since it is impacting me right now, so I appreciate the discussion and would be happy to be wrong.
Eg, Donald Judd’s works are these creative decisions and processes distilled to the most basic of sculptural form.
I'm not even slightly concerned about that.
1. Art is better when it's not paid. Real artists have day jobs that pay the bills and they create art to express their ideas, not to make money.
2. Paid art isn't going away, it will just change. Certain skillsets will be forgotten, like how landscape painting was replaced by photography. But talented artists will leverage AI tools to create works that are greater than anything that came before.
Trying to define who "real" artists are is a folly for the ages. It is the dream of many artists that they get paid for their art, and many achieve it. The starving artist is a mythos of pain and suffering, a good story but hardly good for art. Some of the best composers from history were paid, some of the most influential artists were from wealthy families. They were able to focus on their work without fear of money and because of this they could excel in techique and execution, which allowed them to produce some of the highest forms of their art in history.
Copyright expires, and new artists will create new (copyrightable) art in the future. Unless your assertion is that generative AI is so good no one will make art without it ever again?
This is kind of what happened with music, no? In some countries hard drives, SSDs etc all carry an additional tax that is then given to some copyright organization. Of course it's not the artists that mainly benefit from this, but instead it's the people running said organization.
Eg https://www.telecompaper.com/news/spain-approves-new-digital...
Less equitable world has artists getting paid, in more equitable world, everyone can just use open-source AI tools like Stable Diffusion.
Current market odds for that are at 77%: https://manifold.markets/JeffKaufman/will-the-github-copilot...
If you want to prove your data was used to train an AI, the onus is on you to prove it. Good luck.
The AI who follow the law strictly will be at a disadvantage to those that do not.
which would be easy during a law suit - the process of discovery means you get to check out the training dataset.
The allegation isn't that the AI trainers are hiding, but that what AI trainers are doing _itself_ constitutes copyright violation. AKA, they want the right to use the works to train an ai model to be a right that must be explicitly granted.
i hope that legislation is not introduced to prevent training, as this right would stifle progress.
By the time this reaches judgment and goes through the appeals process there will be a vast industry of non-infringing uses that are clearly transformative and in fair use (Sony v Universal)
You cannot say that the person using ChatGPT to control the lights in their garage is infringing on anyone’s copyright in any manner whatsoever. The point of copyright is not to gain a permanent monopoly on certain speech. The point of copyright is not to make sure that people are fairly compensated for their work. Their work might be terrible but contain a good idea that is later reimagined in a better way (Baker v Seldon) but that’s for the market to decide.
The courts will probably concur that these models are fair-use and I will agree with their judgement.
Yes, and then people say "no" or "pay me". End result of this is that the only ones with good AI models are megacorporations that will DRM the heck out of it.
Years later those same artists will complain that they now have to pay $1000 a year to Disney/MS/Adobe to create art. Because these megacorporations can afford to pay for it. They're the ones that will benefit the most from this, because it creates an insurmountable moat for them.
Copyright exists to encourage the creation of more art and to progress science. AI is clearly a helpful step in that direction. Humans learn from others' works. Should we make that illegal too?
I find it astonishing that people continue to make this argument. A machine is owned by someone, a human is not. Why should the law treat machines the same way as a human? Sounds like some corporate flim-flam to me.
>To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
The purpose of copyright is not to protect the authors, it is to promote the progress of science and art.
The current situation for AI image generation is pretty much the only way these technologies will be available to everyone. Most other paths will simply lead to billion dollar corporations acting as gatekeepers to this technology. Megacorps can afford to hire artists to generate specific art for their AI models, everyone else cannot.
You end up with billion dollar corporations gatekeeping this technology either way (who else has the capital to best train the models?). This isn’t about the little guy.
It shouldn't, that's why arguments that the algorithm is learning, so it's doing the same thing that is legal for humans to do is completely fallacious, on top of it just being anthropomorphism.
The purpose of copyright is to progress science and useful arts. Period. Any action taken in the name of copyright that does not progress science and useful arts is unsupported by law.
What else do we know about copyright? A copyright can apply only to creative expressions. While the bar for sufficient creativity is intentionally low, it is non-zero.
Another thing we know is that purely functional expressions are not copyrightable. When does an expression go beyond being a function expression to a creative expression? That’s up to a judge. Since code is math and math, by itself, cannot be copyrighted, when an expression reaches the level of creative expression must be beyond the math. Updating a database field, factoring primes, or using data correction algorithms are not creative expressions.
Now for AI. Only humans may own copyrights. The output of an AI is not copyrightable. But what if the input was copyrighted?
When it comes to software code, AI will value expressions that are commonly used more so than uncommon ones. But software code is, by it’s very nature, an intertwined collection of copyrightable (creative) and non-copyrightable (functional) expressions. If AI values commonly used expressions, those expressions are highly unlikely to be creative enough for copyright protection in the first place.
So we have a circumstance where AI is trained on copyrighted but Open Source code. Yet the code itself is comprised of both creative (presumably) and functional code, with no clear delineation of what is and what is not protectable.
Lastly, many authors do not understand what constitutes a creative expression that is protectable by copyright. The amount of work required to create the expression is meaningless. Manipulating data to thresh out something interesting is not creative. Let’s just face it that most software is comprised of mostly functional expressions that are not protectable. Back to that “math” problem again!
The big take-away? The purpose of copyright is to progress science and useful arts, not to build walls around ideas and concepts (which, by themselves, are not protectable).
So they can go fuck themselves, or alternatively they can make their super advanced AI reproduce my copyright statement and license every time it copies my code. Which shouldn't be difficult at all.
Is that what you want?
anyways, there's a lot of ways that AI researchers could engage with IP owners to come up with a fair way to use their work, but nobody's making that effort. If my content is part of an AI's training set (and especially if that AI has a tendency to output excerpts of its own training set verbatim, as github's copilot has been shown to do) then it's not unreasonable to set terms and conditions, which could restrict how the content is used for training and what sort of compensation (if any) I deserve.
I'm of the opinion that it's time for new versions of GPL and CC licenses to be created which will enumerate how content can be used for AI training.
I didn't assume anything, just described the likely consequence of your preferences.
Maybe people wouldn't be so angry about an AI trained on mostly open source code if said AI was open source, and not a proprietary SaaS.
Exactly, the point is this one. Open-source doesn't mean liability free, you still have to comply to the license!
There's a reason why some programmers don't even look at proprietary source code leaks as to not accidentally introduce copyright violations into their own code.
It would be. When a human does this, does it invalidate the human's ability to create any new work at all? Should we chain up anyone who violated copyright by perfectly recalling someone's art in memory and re-drawing it from heart, since we cannot trust them to ever create an original work again?
If you copy for your own use only, that's totally fine - or at most a legal greyzone, in the end nobody will care about such personal use copies. If you use AI to generate pretty pictures to hang up in your home, totally fine too.
As soon as you start making money with this stuff though it becomes an actual problem.
It's really as simple as that.
Even the 'generative art aspect' has already been settled long ago when music sampling became popular and required a legal framework.
There’s no reason why we should have the same standards for programs and humans based on metaphors.
If I log in to a website three times a day, I am simply using a website. If a program logs in to a website three thousand times in the span of a second from multiple IP addresses, that’s probably a DOS attempt.
1) SaaS AI people getting richer
2) Devs have less work in the future
I’m not sure if that’s a good development for the open source movement.
But another interpretation is that the generic structure of the code was learned from the works, which is not copywritable. And that generic structure was used to synthesize new code, in much the same way a human who had seen a pattern in a proprietary codebase years ago was able to use that pattern in their own code. I am not a lawyer but most licenses do not prohibit that in my experience. More often in my experience this is what is happening with generative ai.
The tricky bit is that the ai can probably do both in the eyes of copywrite law, since the boundary seems to be very context dependent and existing models don’t have any concept of how much you need to compress and forget the specific details so that it is seen as novel by the courts. The model can memorize significant parts of some inputs despite not having nearly enough space for memorizing the input set, so the first interpretation is possible even if it isn’t the typical output. There isn’t really a kind of “courts will see this as novel” regularizer and there might need to be?
It's not true that that is what humans do.
Having knowledge of where the line is with regards to copyright liability is not an element required to prove liability. i.e. it's of no consequence that the infringer doesn't realize or know that they are infringing. Copyright is strict liability in that sense.
Copyright law has its limitations. But it also has a long history of being applied and interpreted that you can't just wish away. New legal interpretations have to be reconciled with that. So, any radical changes in that are not very likely to happen. Nor will politicians step up and change the law. First of all, few of them actually care or even understand most of this stuff. And secondly, they have plenty of other distractions and a worse attention span than a toddler. Mostly they just do what big companies tell them to do.
So, the most likely outcome here is that these cases won't get very far. At best it might drag on for a few years while nothing really changes. During that time, AI will continue to develop and will get more embedded in society.
And lets face it, this is not Oracle with really deep pockets unleashing an army of lawyers against the likes of Google but some isolated individuals. As legal cost ramps up, their enthusiasm might suffer a bit. Especially if they start losing cases.
This isn't entirely true, due to various entangled trade agreements that require countries to respect each other's intellectual property as a prerequisite.
What the Berne Convention requires is that if I have a copyright in US it will be recognised in France, etc without having to re-register it in every country in the world.
There is no part of the Berne convention that will prevent people from training models on US copyrighted works outside of the US. That is entirely a matter for local jurisdiction.
https://cassels.com/insights/copyright-term-extension-in-can...
And good luck filing a court case under US law in Afhanistan. That's not how trade agreements work. They'd be applying Afhan law (i.e. Sharia law as of recent changes of power). In China it's Chinese law. And in Germany it would be German law. All trade agreements govern is the notion that people need to have some notion of law that applies to things like intellectual property and a way to take legal action when they feel their rights have been infringed. But you get to do that under the local law whatever it is and in the local courts. And of course some countries like China have historically not really done more than pay lip service to such notions. Whereas other countries in the EU have strict laws related to e.g. privacy that don't really apply in the US.
Sometimes we see glimpses of the world that we might have, in the form of open source, sci-hub and independent small sucessess against all odds created by people going against copyright behemoths.
Of course. As an author of OSS, I'm more than happy to let your AI "learn" from my code as long as the trained model is released under a GPL compatible license.
I don't hear complaints about code getting used for training.
I hear complaints about code being used in proprietary products in what appears at first and second glance to be a code laundering scheme without attribution or their rights respected.
If it was trained on the source code for Windows 11, AWS, and Google Search maybe everyone would feel more magnanimous. If those were used I have the feeling that the lawsuits would be much faster.
I don't get how "I'm only sharing this under these conditions" is complicated to grasp. Maybe it's a technical annoyance but... good?
Otherwise every word typed and every image uploaded is contributing to the development of products that will increase the power of mega-corps over time.
That said...
I get it. Huggingface are working on a diffusion based model for music. And guess what? It's extremely important for them to only use opt-in or commissioned training data. Why? Because the music industry, unlike artists on artstation and the likes, have a lot of lawyers and can protect their copyrighted works.
Why should it really be any different for visual arts? I honestly don't see why. That isn't to say I'm not going to keep on using stable diffusion, nor do I think there is anything that can be done to stop it. But, I do think artists should be compensated, and such models should not be based on the work of anyone, unless they want it to be.
> Dance Diffusion is also built on datasets composed entirely of copyright-free and voluntarily provided music and audio samples. Because diffusion models are prone to memorization and overfitting, releasing a model trained on copyrighted data could potentially result in legal issues. In honoring the intellectual property of artists while also complying to the best of their ability with the often strict copyright standards of the music industry, keeping any kind of copyrighted material out of training data was a must.
Reads very different to any statement made by Hugginface or the LION database regarding stable diffusion, which do not mention the concept of an artist or artist work a single time.
And Hugginface are the only ones open about it. Midjourney and Dall-E and Imagen most assuredly doing the same for their black boxes.
https://twitter.com/StabilityAI/status/1605012677188718592?t...
My point probably should be made clearer. You know how you can say "in the art style of X"? Well, it doesn't matter if Huggingface made that possible. You can, relatively easily, train that concept with a collection of paintings by X. Then you can go ahead an make art in their style.
Now, from a technology point of view, that is nothing short of amazing. I still cannot get my head around how absolutely ridiculously powerful it is. And, even the people who play with this, don't seem to fully grasp it either. The world will change in the next 4 years.
What I'm wondering, is, what should we, as a society, find acceptable? Why should someone be able to train "in the style of X" where X is a set of EVERYTHING, and make money of it, without the say of X, or even them getting anything for it? Have you checked the evaluation of Huggingface? It's in the order of 1-10 billion USD.
There is definitely the argument of anti-copyright. I get that part too. But there is definitely someone who will end up with the bigger stick, and it isn't the people holding the paint brushes who made it possible. That seems just a little bit unfair, and perhaps unwise.
Also, I'll end with a point that no one so far has brought up, even though I've followed the discussion both for and against AI. Which is "whitewashing" art. Now, the example I'm going to show isn't very good, but I also spent 5 minutes on this. Where would you draw the line on when Alexander Wild no longer has copyright over his photography?
It's not really fair to take it all out on Stability.AI, as at the very least, they are sharing the technology and models with everyone (for the time being). And that open up some incredible possibilities.
It's much worse what DALL-E, Midjourney and others do, which is much the same, while they let people play with it, but it's all theirs, and they can take it away at any moment.
Or maybe, get this, how about people running AI only feed them information that they legally have the right to use? How is it a bad thing that somebody can't legally steal other peoples' work without their permission because of pesky copyright?
Sure, but you still just cannot output anything that looks like a derived or copied work.
So, maybe ... how about if image generation nets hold onto the training images so that it can compare the generated output against its training data to ensure that it is not too similar.
/s (but only a little)
How much of that is going back to the copyright holders whose work their service derives value from ?
how much of the earnings of the student of art goes to the textbook authors, paintings and learning materials he used to get to where he is today?
But the AI as a whole is capable of reproducing the original in a recognizable form, and it does so on demand quite easily, because it was trained on them - how is it different than selling a zip file containing millions of copyrighted works, and also a bunch of new stuff?
so you're saying that the digits of pi is violating copyright then?
Maybe the AI needs to be able to print out a list of sources to provide attribution. That would be interesting.
that's the responsibility of the user said AI to check.
These AI companies are making serious amounts of money (OpenAI is valued in tens of billions) on the back of artists who never gave permission for their work to be used in this way.
If a child took an artist's work, copied it and made significant amounts of money from selling it then yes they should be within the purview of copyright law.
the copyright aren't all encompassing. There's only an enumerated set of rights granted, and "this way" (aka, training an AI model) is not one of those restricted activities (like distribution or broadcast).
Unless the model can be argued to be a derivative work of the training data set (which i don't believe it is, since the process of training is sufficiently transformative imho), the original copyright holders of the training data do not need to be asked permission.
Your suggestion would be accurate if we lived in a world where we all shared, and there was no money, and copyright didn't exist, but we don't.
But republishing any work as your own, probably falls into that category. And it isn't about profit, but commercial use; thus pasting onto a blog to improve your business (rankings, hit count) is a business use case.
On one extreme:
"Unless you pay your annual Disney fee for having watched Disney films in early childhood, you will need to return your brain to us for processing. Disney was used as the basis for all concepts you know, and as such, Disney owns all subsequent intellectual output of your brain."
And on the other:
In the age of AI, copyright will cease to hold weight. We'll make more new content on a per-month basis than all of recorded human history. The old regime must be thrown away to accommodate the radically new world we're entering.
We'll land somewhere in-between, and I'm hoping it's much closer to (or even precisely) the latter.
I am not fighting for the smaller players but for large enterprises. That is illogical.
That's what they did!
It was in fair use. So yes, they did have the right to legally train the data on copyrighted images.
Many artists don't believe this and the law is very much unclear.
In many cases the AI generated work literally looks like a clone.
is the "right to use in ML training" well defined?
Why should you want a model designed to know all human knowledge to know only public-domain knowledge?
Well, this is the crux of the matter, isn't it? Do you, a human, have the right to look at copyrighted works and learn from them? Do you have the right to use AI to do the same?
- If you bought the book, you can read it.
- If the book is free, you can read it.
- If the painting is in a museum, or on Wikipedia, you can visit it.
- If Bozo the clown says you’re not allowed to look at drawings he posted online, it’s ok. You still can.
Same for AI.
There's a famous Carl Sagan quote: “If you wish to make an apple pie from scratch, you must first invent the universe” which hints at the problem: Nothing is created in a vacuum.
Let's compare what Stable Diffusion does with what Franz von Holzhausen, head of design at Tesla, does. Franz didn't come into existence out of nothing and knew how to design cars. Instead he studied transportation design and worked at Volkswagen, General Motors and Mazda before joining Tesla. All these steps trained his (actual) neural network with inputs from copyrighted car designs.
Based on these inputs he was able to create the designs for the Model S, Model 3, Model X and others. Does this mean that Mazda can now levy a copyright lawsuit against Tesla? It could, based on the reasoning employed in some of the Stable Diffusion suits, but it won't based on the lack of similarity between the cars of both brands.
I believe that the law around AI will come to a similar conclusion. AI learning is neither fundamentally add odds with copyrighted material, nor confirming it. It will be a matter of degrees of derivation - how different is the output of the AI from its copyrighted inputs.
It's also a bit hypocritical, because if you did the same thing to them as a human (in this example let's say be a Tesla copycat) you'd likely be sued into the ground because there's probably a patent somewhere in there.
They take away is contained in the first and last sentences.
Derivation is key.
Copyright protects original expressions, and copying means to reproduce (read and write) something. The analogy OP made is focused on the reading and writing done by humans and the reading and writing done by an algorithm.
Algos like stable diffusion are reading, and their user is controlling what is written.
If the user produces a work that is unique, but uses the style of a particular artist, that seems like it should be valid, since style is simply a process. It's how to create art, but it is not art, and processes are not subject to copyright based on copyright.gov.
With all that said legality isn't morality, and I sympathize with the artist.
So, how do we protect you (as the artist) from this? Copyright, even if flawed, currently protects you from that
I guess my point is more ethical than legal.
Sure, another artist could copy your style, but they would still need to study it for a long time (composition, palette, perspective, grading, line work, etc) to get an accurate understanding of it. They also needed to have spent years training their hand-eye coordination, as well as art theory to achieve it. Whether it's ethical is debatable, but they've earned their skills. And they'd still take a while to produce it, so they earned their money.
If someone without training just asked a computer program to produce "dragon in style of X" then this means people can sidestep artists to create works.
One could be smug and say "work smart not hard" here, but it creates a tricky situation.
What if people start pulling down their works or not posting them for fears a megacorp will crawl their works without permission?
What if someone asks for the same piece (by describing it as a prompt and saying "in the style of X"); is that plagiarism or copytheft?
What will happen to the models in this case? Do they become "inbred" over time?
What becomes of artists? It will no longer be a viable career (especially digital artists) for them. I know most artist enjoy the process more than selling the art, but the process won't put food on the table.
I guess it's a lot of what ifs too here, but they're not unlikely scenarios.
The pure fact that Stable Diffusion tends to produce 3 legged humans shows the complete lack of understanding of its doings.
"true understanding"
I assume you mean the process of looking at an image and not just deriving patterns, but seeing that you are looking at a cat, that a cat is an "animal" which has "four legs and a tail" and that cats can be friendly towards you or aggressive, depending on your own behavior and theirs.
Neural Nets are certainly capable of the first two: classification and creating taxonomies. The last one I admit is tricky as it requires the Neural Net to be an entity within the observed world
"intellectual process"
the intellectual process is arguably exactly the process input->categorize and analyze->compile->produce output loop that we've modelled AI based upon
"creativity"
is the ability to create something truly new. This one seems obvious as Neural Nets only can derive patterns (plus maybe a random input) - but I would posit the question if any human ever created something truly new in the "apple pie from scratch" sense or if we've only ever created higher level works derived from existent things.
"consciousness"
this one is hard to grasp. I would argue that consciousness is the realization that one exists (in the descartian sense) - coupled with the desire to continue to do so. It is a quality that wouldn't make much sense for an output focused neural net like the one behind Stable Diffusion - but it might be a desirable trait in a decision making focused deep learning setup - similar to a self-healing cloud deployment.
"love/emotion"
This builds on the previous consciousness example. Not to sound like Rick Sanchez /some other cynic - but aren't these at their core adjustment mechanisms that help us further evolutionary goals like survival and continuation of our lineage. Wouldn't a decision making focused deep learning setup be more stable/have a higher uptime if it would facilitate its goal of "staying on" through a strong drive of survival/expansion?
The last two examples are where my point falls apart a bit. But I still stand by my general thesis: We are way too certain that our particular human way of processing information and "thinking" has some divine quality to it that isn't replicable in neural networks. Against that, I would argue that neural networks are largely the same mechanism we employ in our thinking and that they are just a couple of millenia in evolution behind, but are catching up at a multiple of the speed it took us to get to where we are now intellectually.
Keep in mind, we don't know how humans learn either (on a neurological level). It might end up being that we stumbled onto the same general idea, using matrices and linear algebra instead of neurones, synapses and neurotransmitters.
See also the story of Trurl's Electropoet (from the Cyberiad). Trurl first had to simulate a universe to get the poet to work.
the cat is out of the bag, the worms are out of the can, the feathers have blown away in the wind
these developments seem very likely to be central to programming, all other kinds of engineering, conceptual art, scenography, costuming, technical illustration, and pornography, within a couple of years, even if (against all odds) development on the neural nets themselves makes no further progress; they enable you to do things in minutes that previously would have taken days, things which are core parts of the feedback loop driving these disciplines
if every country in the world except thailand bans it then within ten years all your kids will be secretly watching prohibited thai movies with software secretly written in thailand on surreptitiously thai-designed computers, riding thai bicycles
even if deepfakes mean that the most significant effect of ai art is enabling massive fraud, spam, and mitm attacks, banning it locally won't stop you from falling victim to it (fraud is already illegal) but just from developing effective defenses against it
Furthermore, if we adopt current copyright laws to AI rather than understand the entire world is changing, only the largest AI companies will be able to leverage the technology to train their models.
If it requires every film in existence to train a model, only Disney or Disney licensors will be able to operate. That's not good for competition. It might make it even harder to grow up as an independent creative or startup as it presents an impenetrable moat.
As I currently see it, weakening copyright is the only way to assure democratic access to this technology.
Sure, existing works already licensed can still be used, but at least both parties (copyright holders and AI trainers) won't have anything to argue about.
[1] Anyone from CC reading this? Make it the default.
The moment you share your creation/work to someone/the world, you are training their nn.
You can not share something publicly and then demand "xyz" can not view it. Viewing is training.
You are free to keep your creation under lock & key and only share with nn (of people and/or AI) of your choice.
That's nonsense. Licenses have clauses on how the content may be used. Clauses along the lines of "The content may not be used for ..." are common.
I dunno where you heard that once you release something the license clauses no longer apply, but it's wrong.
That's news to Microsoft[1], who's shared source and various NDA licenses for the source code already has clauses restricting what you can do with it.
[1] I think the problem is that the pro-AI arguments are coming from people who are not aware that clauses in licenses restricting how the content is used is quite common. For example, you were obviously not aware that they were so common that almost every big tech and/or software company of the past and the present already have those clauses in, and those clauses have already been found to be enforceable!
unfortunately you have descended from simply making vaguely ignorant comments to attacking me, which indicates that further engagement with you is unlikely to be useful to anyone
And every nonFLOSS license and I've seen has restrictions on what the licensee can do with the material.
Can you find one nonFLOSS license that doesn't have restrictions or limitations?
I mean, you lead with obnoxious, then descended into condenscension, all while not realising that the whole point of the license is to restrict the licensee.
That's a different problem. Let's not get into the argument of "Just because the victim cannot prove something, we should remove the relevant laws."
The current laws are sufficient; all that it takes is for licenses to have a non-AI-training clause.
Lets solve the problem of "how do you prove" when we get to it[1].
[1] Right now, due to the systems already trained being given every single image on the net as training data according to the owner of those systems, in a civil suit the burden will be on them to prove that, on the balance of probabilities, a particular image was not used.
granting many such licenses is a recipe for social collapse
Who said they couldn't be enforced? The argument from the pro-AI team is that we shouldn't have those laws in the first place.
I'm saying, let's keep the laws we already have because they already work quite nicely if the content creators don't want their work included in any training data set.
Rushing to make new laws because "Muh AI" is silly.
if you are failing to understand the factual claims, you have no hope of making sense of the normative claims for which you have such thirst
Almost all big companies routinely have licenses which heavily restrict how their software may be used, and the licenses have held up time and time again in various jurisdictions around the world.
The only argument you appear to be presenting is "well and don't like it that way".
Tough. It's already that way and has been so for dozens of decades .
copyright is not a get-into-jail-free card that allows private parties to invent their own legal system and nonconsensually impose it on other private parties
Forbidding certain uses is a particularly common clause in most copyrighted material.
Other than some FLOSS licenses, can you find one that doesn't have limitations on what you can do with the material?
With generative AI I'm mostly concerned with what it will do to the next generation of artists. I don't think I would have ever had the motivation to pursue music if it had been possible to replicate perfectly with AI. I'm immensely thankful I got that opportunity and so I want the next generation to get it as well.
i think the second paragraph is a bit myopic, like hunter-gatherers observing agriculture and worrying whether the next generation will be able to track prey through plowed fields, or will allow their hunting skills to decay because it's easier to get meat by trading with the agriculturalists
but that understates the case; ai (if this is really ai this time) is certainly a more significant innovation than agriculture, probably more significant than fire, on par with tools and language
still, that says nothing about its moral valence
afaik asahi v. superior court is still governing precedent in the usa though so it won't be of any interest to domestic litigators in the usa
Extraterritorial jurisdiction?
You sell stuff to country A, you comply with laws of country A. Which is why USA companies have to take GRPR into account.
You make an illegal model for country A? Can't sell it there.
Well they allow mining the data but nothing is said about the copyright of the collage output.
occasionally outputs training data verbatim: yes
verbatim output is somehow not a copyright violation: ???
It reads like there's a bunch of countries that have similar legislation, interesting though.
This is too reductive.
The back pressure ML is generating is at this point too strong for anything to make any difference.
This is wrong, regulation can make a difference.
Anyone who attempts to stop it will just be practicing self-sabotage.
This is a prospect worth evaluating organically. Learning the potential is much different from accepting self-fulfilling prophecies.
Regulation is per country or bloc. With ML the value of defecting is so high that any regulation you impose on it which restricts its utility will amount to self-sabotage.
but i agree that in this case it's probably futile
It’s a pretty simple problem. The folks claiming it is not either are being intentionally disingenuous or honestly do not understand the legal definition of derivative works in copyright. It’s settled law.
You could… you know either a) don’t train on works you don’t have a license to or b) use some sort of adversarial training to ensure that the AI doesn’t replicate the work it is trained on.
I could see this go either way. There's the argument you put forth, and then there's the argument that a text to image model is a transformative work. You can use copyrighted works and make money off your product and still have a transformative work. The Google books case is, of course, good reading on the subject.
My main point is that it is not at all clear which way the law will go on this.
This early paragraph is so bone-headed, so smugly demonizing, and misrepresents the situation so badly that I had to stop reading. This paragraph isn’t analysis, it’s propaganda.
What a gross and hateful way to frame this, especially given your apparently limited experience…
If it’s a copyright matter, I don’t see how that could work. It’d need to be opt-in, or covered by an explicit license grant (and terms and conditions are being ignored to the point that I gather some jurisdictions’ courts are pretty much striking down anything that a reasonable person wouldn’t expect to be there, and so it wouldn’t surprise me if such an approach would strike down any grant asserted in T&Cs).
> for software authors, prohibiting ML training would be antithetical to the Open Source Definition
If it’s a copyright matter, this wouldn’t be the case, because it wouldn’t be discrimination against a particular field of endeavour, but rather simply insisting on the terms of the license.
It doesn’t solve attribution, and you could get a sufficiently advanced ai to “launder” copyright but at least it would prevent those corporations from leveraging public works into copyrighted projects
It'd stop buisnesses from replacing artists with models wholesale, but it would allow people to keep using them, and allow the companies making the models to make money selling access to the models.
That is legally the default. Creators own their copyrights. In many cases it is made explicit with a creative commons non-commercial use license. Remember, without a license you get nothing commercial - except the nebulous fair use.
The real problem here is companies thinking they can consume large amounts of material and works simply because they can see them on the internet and obfuscate them by combining together.
Just because AI art models can reproduce copyright (if you try hard enough), doesn't mean that it's a copyright stealing machine.
It's the responsibility of the maker for the AI model. They train the AI model on copyrighted material. So they should clear the licensing for using that material.
IANAL, but since I also work in the ML domain, I tried to find out how this works when you have to follow EU laws. Past rulings [1] have considered 11-word snippets to be a copyright violation (which is by no means the lower bound). So, it is likely that if a copyright holder in the EU can show Co-Pilot or ChatGPT to reproduce a non-trivial fragment of code or sentence, that a copyright holder can sue them successfully.
However, the sad fact is that these models are made by well-funded entities. So, they'll bury small copyright owners in lawyer busywork until they go bankrupt and settle with big copyright holders. So one possible net outcome will be that large entities can do large-scale copyright violation while individuals and small companies can't. We have seen this story before. And it helps to entrench big companies even more.
I hope that the EU comes with some regulation to level the playing field. So either make it illegal for everyone (enforced by EU courts) or legal for everyone.
[1] https://www.theregister.com/2009/07/31/ecj_rules_11_word_sni...
* * * * *
[0] https://sso.agc.gov.sg/Act/CA2021?WholeDoc=1&ProvIds=P15-#pr...
About time.
> If the AI industry is to survive, we need a clear legal rule that neural networks, and the outputs they produce, are not presumed to be copies of the data used to train them.
But they are compressed lossy copies of all that data! That's the whole point of noise/denoise functions that neural networks are based upon. The whole mathematical foundation of training a neural network is "teaching" it how to recognize and/or create copies of data stored in the training set.
> Otherwise, the entire industry will be plagued with lawsuits that will stifle innovation and only enrich plaintiff’s lawyers.
When you're willingly breaking already established law en masse for profit in hope no one cares enough, be it copyright law or any other, you're not an "innovator", you're a criminal. The fact that you're a tech giant or a Bay startup doesn't matter in this regard; the only thing that matters is the notable amount of time required for the justice system to catch up with your novel tools for laundering intellectual property.
If anything, your comment highlights the need to evolve IP norms and laws.
With NNs trained on thousands or millions of data entries, this concept becomes fuzzy in the same way as you described - a short summary likely wouldn't be considered a copy, just like a 64x64 generated thumbnail wouldn't be considered in the same way a 4096x4096 hi-res image.
I haven't seen that happening since the discussion started. Most of the complains I saw aimed at things like "it stole my style" not "it reproduced my art".
Do you have any examples?
In music you aren't allowed to use the same notes, even if you played them on a trumpet with a swing beat, while the source was on the piano very staccato.
While we don't have the same vocabulary for art, it's not unreasonable to expect similar protections.
https://www.vice.com/en/article/wxepzw/musicians-algorithmic...
Your example is one where nearly no work was done, thus it doesn't deserve much value. "Let a = the set of all songs" doesn't help me find new songs I like. A songwriter does that work. Another artist that takes and uses and resells that work (without consent), is stealing that work.
To me it's funny that nearly all the problem with the current team of AI generation would be solved if the model generators simply licensed the content they train on. "But that would cost too much" Ok, just use public domain work, "But that wouldn't be as good" Oh so you are saying the work has value, but you are unwilling to pay for it, and instead your scheme is to just take it. That seems like a good definition of stealing - not paying for something that has value.
Another artist accidentally uses a melody from another song (because it's a finite set) and are sued for all their income is a horrible system. The winners aren't the people producing value, it's the people who got there first and are now profiting off other people's work.
This is so common the recording industry itself has established rules for sampling and licensing and covers and what not. Are there some folks out there abusing the system, for sure. But overall its goal is to maximize the value produced by the recording industry, which very much includes the people who 'got their first' who built foundations for future artists. To me, this all seems basically reasonable.
You are aware that there is very expensive art out there where the artist did not much work. Like painting a canvas in one colour or throwing an item in the corner of a museum.
According to you, that would not deserve much value but it does have a lot value in reality.
In fact "value" is what somebody else gives to the piece of art.
A prompted AI artwork made by me may have more value to me than all the art in the Louvre.
The discussion here continues to turn around copies when it's not a copy those algorithms generate.
Pretty much none of these systems "reconstruct an image in detail".
[1]: https://arxiv.org/abs/2212.03860
The fact is that these systems are complex, new, and interesting. However, it is not the fault of small-time programmers and artists that modern copyright law is a major, overreaching mess that is now finally greatly affecting what the big corporations want to do. They are getting sued? Cry me a river… Perhaps they will finally stop backing the American-led copyright lobby then?
From a quick skim of this paper, they apparently used toy models with a few hundred to a few thousand images in the training set. For the ones with as few as a few thousand training images, they rarely or never saw exact duplicates.
For instance, in their figure 4, they show exact duplicates for the training set with only 300 images (well, duh), and didn't find any exact duplicates for the training set with only 3,000.
I'm not sure I'd call this a "strong argument" when applied to models with millions or billions of images. Quite the contrary. LAION-5B (used in Stable Diffusion) was trained on 5 billion image/caption pairs.
They explore a range of sizes and I do not think it is fair to to only highlight the smallest ones. They do explore a 12M subset of LAION in Section 7 for a model that was trained on 2B images. Yes, it is not an ideal experimental setup to use a subset (they admit this) and far from LAION-5B, but it is a fair stab at this kind of analysis and is likely to lead to further explorations.
Let us return though to your claim, which is what I objected to: “Pretty much none of these systems ‘reconstruct an image in detail’.” I think it is fair to say that this work certainly makes me doubt whether none of these systems (even the larger ones) exhibit behaviour that may limit their generalisability or cross the boundary of what is legally considered derivative work.
You may very well be right that once we scale to billions of images this behaviour is improved (or maybe even disappears), but to the best of my knowledge we do not know if this is the case and we do not know when, how, and why it occurs if it does occur. I remain a firm believer that these kinds of models are the future as there is little evidence that we have reached their limits, but I will continue to caution anyone that talks in absolutes until there is solid evidence to support those claims.
For example, I could take a massive 8k video and covert it into a very small 144p youtube video. Am I in the clear simply because the output is tiny compared to the input? Similar I could take a huge studio master copy of a song and convert it to a very small and rather compressed (distorted) mp3.
I partially agree that some of the problem is when perfect copies are spit out by the models, but I do think there is a bigger problem. Copyright is a complex concept that can't be defined exclusively by a single metric like size, and any mathematically definition will in the end be killed if large copyright holders feel threatened by it.
"Transformative Use" is a major consideration in fair use copyright: https://en.wikipedia.org/wiki/Transformative_use
ML models do not supplant the pre-existing work, and provide fundamentally new modalities. Transformative use seems like a slam dunk to me, but I guess we'll see what the Supremes decide in twenty years or so...
"Some courts have held this factor to be the most important in the analysis."
But when it comes to market harm, does the tone of my review effect the enforceability of copyright?
As in, if my review is negative it would harm the market for people going to watch the movie vs a positive review right?
A review can be commercial, can cause significant harm to the market, can include substantial amount of the work, and yet the character of use can be significant enough to convince a judge that a exemption should be applied. Since judges historically has come to this conclusion there exist now legal precedence. With precedence we can make some general conclusions which tell us that reviews are in general exempted when using other peoples copyrighted work for the purpose of reviews.
This character of use is very different then if I convert a studio record of a song into mp3 and publish it on p2p sharing site. Judges has historically viewed the character of use in those situation as not being worth giving exemptions.
You're not directly competing with the movie though, your work is a review, not a feature film.
If you were to make a parody movie from the material of the movie itself, directly taking scenes and altering them to your liking but still relying on the viewer recognizing the original in it, you'd have a harder time, I think.
[1] https://petapixel.com/2023/01/17/getty-images-is-suing-ai-im...
> So, all neural network developers, get ready for the lawyers, because they are coming to get you.
This is just dumb.
I happened to have missed those discussions, do you have some links you can point towards? thanks!
https://githubcopilotlitigation.com/
Both along with their respective HN threads.
Even in my overfitted dreambooth model of my wife it doesn’t pop out the exact same portraits.
I have seen very few practical cases of generative models such as copilot and even less of a stable diffusion of reproducing original copyrighted works in exact detail, and the few that I did encounter were instructed torturously to do so, which strikes me as highly contrived.
This feels like an argument for communism being the most productive system in theory. Most of the time I feel like I see uninspired material that’s tracing it’s own training data.
Machines currently are adept at making copies within a delta, hence articles such as this to limit copyright so the people who operate said machines can profit.
So are those you are storing in your mind...
(Edit: i.e., "learning is not a violation". Also see below.)
I do not see much space for misunderstanding: which relevant box takes instances of input data to output one of them?
Edit:
Make your point explicit, sniper... You can hide in front of the ### consuetude of "silent disagreement", but it remains violently annoying. The article contains "nuances" like "derivative work", but the poster is objecting to the article calling «neural networks» not «copies», as «compressed lossy copies», and I retorted that so is everything you learn, and holding them in the "corpus" is not considered "retaining a copy". If you have objections to that, either you present them, or there is no contribution in shaking your invisible head.
Drawing Coca Cola's logo from memory by hand and slapping it on your product is still copyright infringement, even if it isn't an image Coca Cola has ever produced. In that sense it doesn't matter at all if it's AI or human - the production and subsequent distribution for profit of a copyrighted thing is not allowed, period.
The current set of AIs do exactly this all the time. That's a very clear legal problem.
I think the point still stands though. Just mentally replace it with "a random DeviantArt" and it all still applies.
Sounds similar to, for instance, songs written by people who have heard other songs. I wouldn’t expect legal cases concerning AI-generated works to be any simpler than legal cases concerning the difference between a songwriter violating the IP of another songwriter or simply being inspired by another song.
~~robots.txt~~ ai.txt for code repos? :P
It's all bad. They really should pay the "little man" not just publishers with lawyer budgets - same with AI.
If you're a designer and you take "billboards or t-shirts or anything in the real world that is copyrighted" as stored in your head as the basis for something you're working on, you will need to consider ways in which your derivative work may be infringing.
Similarly, no, the space where copyright meets training data doesn't "means Waymo, Bing, Google are all illegal." Nobody is going to care if a driving or search neural net has that billboard or t-shirt data, because their function isn't to output copies or derivative works.
If the function of what you're building is to output billboards or t-shirts or anything in the real world that's copyrighted, then you may violating the spirit of copyright law whether you're wetware or using silicon. And whether or not the letter of the law has been refined carefully to apply to the issues at hand.
The creators of stable diffusion didn't create it with the intent helping people infringe on copyrights either. But some people will use it for that purpose.
You cannot then use this copy commercially while pretending this indirection somehow grants you immunity from copyright laws.
The fact that AI is involved does not change the basic principle.
It shows that a lot of people in this space only cared about standing up against corporations, they didn’t care about the philosophy behind the anti-IP movement,
I think if we tried to calculate the cost to humanity of the things that don’t happen or aren’t created, are very expensive to do, or are restricted to people and companies with the right “rights” we’d uncover a tremendous tragedy.
In addition to the first order effects (you can’t do X or have to pay to do Y), there are huge chilling effects on uses that are allowed (especially as companies regularly overplay their hand well beyond the protections they are actually afforded), and the cost of technology and people that exist only to enforce these rules.
We should embrace the free sharing of information and maximisation of its value for all of humanity.
If we worry about researchers, programmers, and artists not getting paid (and we should!) then we ought to pay them as a society for the public goods they create (as we partly do for science and research). Finding and implementing a good and democratically reasonable way to do this would be truly revolutionary.
I hope that AI research and training only goes to prove the long term damage and futility of IP as a concept and accelerate its downfall.
I am, a wee bit, shocked people believe there can be impactful legislation on this. As if politicians who have been unable to curb PIRACY in any real sense would now be able to tackle an even tougher problem. This is despite well funded lobbying groups. Even large corps enable piracy without consequences.
Further, the government frequently indicates a concern that China will beat the US at AI.
There is an extremely bumpy ride coming for a group of people that have never had to deal with an unavoidable bumpy ride. I look forward to an increasingly logical viewpoint from people being struck by reality. Not maliciously, but societally.
AI has indeed removed the "need" for copyright. Let Mickey die.
If its an image editor everyone is calm, but call it AI everyone loses their minds.
Like it or not, you can’t stop people from copying strings of ones and zeros. Time to let go of primitive notions of copyright and embrace that we’re all one species and we benefit by sharing our knowledge.
That's why the "forgiveness instead of permission" approach is seen as somehow heroic, when it actually simply is abuse.
I'm not sure why the author is out here talking like a lawsuit brought by the author of Typography For Lawyers is going to bring down Microsoft like it's a foregone conclusion, but people are going to be training models on public data for personal gain from here out no matter what happens. The cat is out of the bag.
If copyright ends up being enforced, it’ll affect commercial deployment. We’re in for an inevitable paradigm shift soon anyway, so if we need any guardrails, we better install them before we open the floodgates.
I don't really see why humans are treated different than a computer, for this case. A human also learns from lots of copyrighted material. It's not possible that whatever a human has seen or has heard, will have no influence whatsoever on the human brain. So by the argument here, everything what a human does, ever, is always derived work from everything he/she has ever seen in his life.
So, then another argument is, of course there needs to be some line. It's only derived work or breaks copyright if it is really similar enough. But if this is now the argument, where is the problem? We can just apply the same to the AI. Of course, this is somewhat ambiguous, where to draw the line, but it's just the same as for humans.
Another argument is, the AI has in total seen much more visual data or text, than any human ever could in his/her life, so that is how the AI is in any case different. But I don't really see, why is this relevant? Some humans are reading more books than others. So those who have read more books are in danger, at some point to have read too much books? Where is that line?
Another argument is, stochastic gradient descent works different than the human brain learning algorithm. I don't really see how the details of these technical difference are relevant here.
Another argument is, the human learning is much more efficient in terms of data. But I don't understand how this is relevant here. Isn't this actually an argument in favor of the AI regarding this topic?
Future research on AI might make the AI, its behavior and its learning, more similar to humans. But if we now have a law which says it cannot use public copyrighted data to learn, then the AI has a huge disadvantage to humans, because humans use such data all the time to learn.
This is the exact same progress that music sampling went through.
So, all neural network developers, get ready for the lawyers, because they are coming to get you.
No, you dullard child. Get ready to get sued if you try to make billions of dollars via derivative works of other creators while breaking software licenses.This really shows you don't know the movement it self. People want credit, and sometimes put conditions to use their work(GPL and copyleft), and when the AI doesn't follow these guidelines, then its breaking these copyright laws.
Not everyone is willing to willy-nilly give their code for nothing.
I didn't expect to be defending copyright law, but the excerpt is ridiculous. It's clear that the images are a product of the prompt and training set. The legal side of fair use and copyrights is best left to the courts, but the real question is how to divide the profits. Because the training dataset is so large and the technology plays such a crucial role, it may not make sense to pay much (if anything) for each individual image in the training set. There's no clear cut answer.
My only hope is that somehow generative art, code, and all other media can somehow benefit the little guy rather than continuing to enrich the largest and most powerful, however unlikely it seems.
Thanks to lawyers, a technology that had the promise to democratize art will be used by large corporations to enslave us further.
Worst case scenario - you're right, they're willing to go through that and the authors actually get paid something rather than nothing.
What a disingenuous framing of the open-source software movement.
When I license a work I created as GPL or MIT or whatever, I do so because I imagine that some user or some company would use the software to build something based on my software and contribute their changes back to the community so that we can all benefit. This is standing on the shoulders of giants. Microsoft "using" my source code to build a for-profit programming-as-a-service bot was not what I had in mind. At the very least, the complaint that the model doesn't give credit to its sources is correct. We might need new licenses so people who write open source code can opt out of this type of usage.
> The Stable Diffusion suit alleges copyright infringement, stating that, “The resulting image is necessarily a derivative work, because it is generated exclusively from a combination of the conditioning data and the latent images, all of which are copies of copyrighted images. It is, in short, a 21st-century collage tool.” That characterization is the essence and conclusion of the lawsuit, and one with which many AI designers would disagree.
AI designers can disagree, but how about they try building an AI model without using someone else's images. The AI model is not even possible without the source images. Show a little gratitude.
If you ask me, AI is not in danger of being swallowed by copyright law—it's in danger of being swallowed up by its own self-entitlements.
https://www.nolo.com/legal-encyclopedia/fair-use-what-transf...
And then there's the legal aspect: in music, sampling is legal nowadays, but you have to ask the original author for permission and then pay royalties based on how much your own creation is based on the original sample.
AI isn't much more than automated sampling, how can AI generated content ever hope to avoid a legal quagmire if it cannot properly attribute its source material?
Will the person be able to monetize that video.
Say a person makes a program which automatically makes a collage of movies with some music on top. The movies will be credited in the video description.
An AI is also a program. First of all, why can’t it credit its inputs? And second of all, should it be allowed to facilitate monetization on behalf of whoever owns the program?
It seems to me that if the "LLM Copilots" just observed the existing license there would be less of a problem here.
Copilot, only recommend work based on $LICENSE or $LICENSE compatible license when I am working on $LICENSE code.
What's the problem genius?
AI is just a tool, do you also sue the company that sold the paintbrush with which an infringing painting was made?
However they don't give you any way of NOT infringing when using their tool.
What we are looking at is a kin to a company that makes paint brushes, trains graphics artists, and contracts out their graphics artists to use those paint brushes. If it turned out that the graphics artists turned out derivative works without the rights holder's permission, you can be assured that people would want to "sue the company that sold the paintbrush".
I'm not going to claim that my example is equivalent to what is happening with these AI services. And while you may be right about AI fundamentally being a tool, like a paint brush, I would suggest that is only true if you ignore the data that is fed into it.
Microsoft should also provide resources to content creators to find where their work has been used.
This is not complicated.
Generally I'm reluctant to say people are acting in bad faith. It's less tough to say Microsoft is acting in bad faith here.
My point here is Microsoft knows there's no way to tell where the suggestions are coming from, they stripped out that information, in this case it's even more likely that the offense happens because they're selling the paintbrush on a large scale. It's just a question of how often it's going to occur. Is it for 30% of your users or 70% of their users that are infring on licenses on a wide scale basis?
So you should know and list all of the authors that contributed to that code being generated.
And that's the least restrictive software license…
"AI is in danger of being swallowed up by copyright law" is perhaps the oddest take on the current situation: people's copyrighted work being in danger of being swallowed up by AI.
Does anyone really thing this anti-Copilot case has a chance of winning when my guess is that Copilot adoption is exponential over time.
Strongest counter argument is the Supreme Court struck down abortion even though abortion was rising in popularity. So that's a bit worrisome.
True, but in order to train these models you have to copy and save images that are for the most part copyrighted. So while the models are not copies of images, they are definitely derivative works.
"Let us train our LLM on your content in order to prioritize your content in search results"
Obviously, creative people have always stolen from others. But it required a lot of time to immerse yourself in the work of others. With AI, it’s only a push of a button.
If everyone is somehow this r-d, there's China also.
Rather than settle this ambiguity on court, why not remove it altogether? Personally I think we should be adding new clauses to licences that either explicitly approve or prohibit the use of code for training of models.
Sheesh.
Now looking at the current state of AI models, they have shown to reproduce the training material in large verbatim pieces. On the other side, they lack any ability to reason about what they produce. Which for me is a big indicator, that they are clearly reproducing not producing on their own. Which creates the potential copyright issues. At minimum, the legality needs to be clearly specified, be it by court rulings or law changes.
Ironically, to some extend this has been a long time problem with human creations too. There are plenty of lawsuits about music pieces "borrowing" too much from older creations, and recently there is a growing cases uncovered, where expecial thesis papers contain too much unattributed content. Probably no new thing, just so much more easily discoverable thanks to modern text search tools.
So as a tl/dr: I don't think current AI models are comparable to human learning and the copyright questions around that need to be decided. If that means generally expanding fair use, having shorter copyright periods, I am all for it. But "AI" shouldn't be a tool for large tech companies to systematically evade license terms.
but, to me, there is a big flaw in these lawsuits. if we need to shut down these AI technologies, then our entire education field is in danger. the same arguments should apply to our education field.
educating AI is the same as educating ourselves.
Sounds like a courtroom is the right place to find an answer? That is the the US system isn't it?
Serious naive question: How would that not be a desirable outcome for society? “AI” is becoming a scourge.
AI’s value-add that we truly need is just detecting cancer right? Obviously I don’t want anyone to die from cancer if they could be saved. But AI in cancer detection is a detection-rate booster: cancer can be detected without it. A non-AI process could try to recover some of the lost accuracy.
Weigh this against the misinformation potential of chatGPT, deepfake video fraud, discriminatory bias in ML model output, the surveillance potential of image recognition, addictive social media, the essential inscrutability of model output… (i could go on, but the research & reporting on AI as we have wielded it is voluminous)
So if litigation kills AI, isn’t that cause for celebration, on-balance?
Or shit, can’t we just legislate easier usage regulation for lifesaving medical data, so we can keep the cancer detection use cases and let the rest of the AI gold rush die?
That’s a big sad. Unfortunately that’s called the real world. You also can’t break into someone’s home and film a movie in their house without permission. That doesn’t mean movies are illegal.
I think that the solution is easier than it seems. The real value websites like Getty provide is not the image in itself, but the labelling. Put the labels under a paywall and problem solved.
For source code it's more complicated, but I guess that you could somewhat limit access to commit comments.
In any case, the genie is out of the bottle.
Exactly. Same back then as today.
An insecure cave person would be protective of their cave painting turf. A well-adjusted cave person (also likely more talented) would be sharing ideas and resources or have other work for/with them.
the concept is very new
Really? Curious that we know the author's name then… seems like… attribution?
Do you get angry if someone "steals" your programming techniques and ideas, and copies them into a new portion of the codebase?
I think all of this is very cultural.
If your AI is spitting derivative works or straight-up verbatim copies of the original, stripped of a license, and given the original license even allows derivative works, then yes, you better lawyer up. This applies to both software, art, or any copyright-protected work. Legality aside, it is also shitty and on the lower end of morals to abuse people's work/art in this way. AI systems that don't fuck with people and their work are perfectly fine. Stop fucking with people and you won't need lawyers.
So then I guess these new AI models are only problematic if someone tries to sell them (as its essentially selling copyrighted works) but an open source model is probably fair game. Individuals can use it, but if you try to monetize you better not be infringing.
Seems like something like YouTube’s content ID solution for detecting copyrighted works could be a “quick” fix here. You can sell access to your model but not to generate copyrighted works
What do you think Microsoft is doing with Copilot and everyone's code on Github? They didn't spend millions of dollars and several weeks of training just for the amusement.
Pay people whose works you want to extract value from, it is that easy… Unless the whole point of this "AI" bubble is to create monetary value for the companies and their shareholders without being held accountable. I do not mean that creating, configuring, tuning a model, or even compiling a massive dataset is not work, but it is the tiniest fraction of the work that went into whatever is present in the dataset.
> also, for software authors, prohibiting ML training would be antithetical to the Open Source Definition. So that probably won’t work.
A-ha, at one point I hope even AI zealots will be forced to acknowledge that the process is creating a derivative work and sure, train on my FOSS sources all you want, but the end result will need to abide by the licenses of all the sources browsed (have fun).
> As the tech industry celebrates the frothy emergence of machine learning in a time of economic doom and gloom, let’s hope this nascent field doesn’t sink because of the copyright iceberg looming ahead.
I am sorry to break this to the author, but not all fields and jobs need to exist. I have not seen much net positive from "AI" to society as a whole so far, with even less exciting things on the horizon.
So, e.g., from Basic Attention Token (BAT) to Basic Authorship Token (also BAT)?
People keep pretending that AI is the exact same thing as a human learning, but it's really a lot more like a fancy compiler with highly non-deterministic output. AI is not a human.
https://www.abc.net.au/news/science/2023-01-12/chatgpt-gener...
This is disingenuous... the issue for FOSS is the scrubbing the license off the code and then users not observing the terms copilot or whatever concealed from you... that is against copyright law.
It's not a real problem? Great dump MS monorepo in there for public use.
Especially strange coming from a site which has 'copyleft' in its name.
Here's a little(?)-known fact: the leaked Windows XP source code is hosted right on GitHub. Has anyone tried typing some win32 API function prototypes and seeing if Copilot fills in the function bodies? I'm not a Copilot user or I'd try it myself.
There's a Windows monorepo, and a monorepo for chunks of Office, but even now that many teams are on Git they're still very much different repos with different build tools, different code search (maybe they fixed some of that in the last few years but none of my friends there have mentioned it if so), different coding standards, and more.
That's mostly what folks are referring to. We can't really take Microsoft at their word about systems like copilot if they're unwilling to trust it with their crown jewel.
But please, please, protect art making from machines. These paints made in those caves werent done during work hours, it was probably the first forms of leasure our ancestors experienced. The first forms of enjoyment in the rude life of early humans. I think there is some higher order aspect to art that's related to our essential well being.
All Im saying is protect it; im not saying AI shouldnt be used. It could certainly render art making even more fun. Maybe limit the resolution of generated pictures for example.
There will be innovation outside of the latent space thanks to the AI art movement, that will be a positive development. An enormous body of artworks has zero online presence and exists outside of its reach, at least for now. However, art reflects its time, and generative, machine learning, text to image art will all influence artists in unseen ways. Offline artists are watching.
Nobody needs a computer to do outstanding art. Even the most primitive tools can produce unmatched results in the real world. That's the beauty of making art. Models cannot even begin to touch it. Good artists have nothing to fear about ML as a tool, which is fantastic and promising as such. It's undoubtly diruptive. As image-making takes less time, the value of an individual image will tend to zero. Prices will go down, and artists will adapt. That's what they do. Sure, there will be a bunch of profeteers and imposters using AI. Art always had its share of those. Art is a dream field for fraudsters. A known issue since we commercialised art, because it carries intangible value. Hence why NFTs started with art. That doesn't make the whole initiative a fraud. That said, digital artists who have copyright issues should absolutely fight back. The gray areas around training model with art by living artists needs an open discussion.
The line you're drawing is entirely meaningless.
This reads like an argument from the English Luddites[0]
0. https://en.wikipedia.org/wiki/Luddite#Birth_of_the_movement
However, what’s happening here is different. What people are reacting against is that (often dirt poor) artists unwillingly become part of a commercial supply chain, while corporations protect their IP the usual way, backed by the full force of the government even on foreign land.
If copyright is abolished entirely, at least there’s some form of level playing field. But this becomes yet another step towards a more exploitative economic environment.