Or maybe, get this, how about people running AI only feed them information that they legally have the right to use? How is it a bad thing that somebody can't legally steal other peoples' work without their permission because of pesky copyright?
Or maybe, get this, how about people running AI only feed them information that they legally have the right to use? How is it a bad thing that somebody can't legally steal other peoples' work without their permission because of pesky copyright?
Maybe the AI needs to be able to print out a list of sources to provide attribution. That would be interesting.
that's the responsibility of the user said AI to check.
These AI companies are making serious amounts of money (OpenAI is valued in tens of billions) on the back of artists who never gave permission for their work to be used in this way.
If a child took an artist's work, copied it and made significant amounts of money from selling it then yes they should be within the purview of copyright law.
the copyright aren't all encompassing. There's only an enumerated set of rights granted, and "this way" (aka, training an AI model) is not one of those restricted activities (like distribution or broadcast).
Unless the model can be argued to be a derivative work of the training data set (which i don't believe it is, since the process of training is sufficiently transformative imho), the original copyright holders of the training data do not need to be asked permission.
Your suggestion would be accurate if we lived in a world where we all shared, and there was no money, and copyright didn't exist, but we don't.
But republishing any work as your own, probably falls into that category. And it isn't about profit, but commercial use; thus pasting onto a blog to improve your business (rankings, hit count) is a business use case.
is the "right to use in ML training" well defined?
Why should you want a model designed to know all human knowledge to know only public-domain knowledge?
On one extreme:
"Unless you pay your annual Disney fee for having watched Disney films in early childhood, you will need to return your brain to us for processing. Disney was used as the basis for all concepts you know, and as such, Disney owns all subsequent intellectual output of your brain."
And on the other:
In the age of AI, copyright will cease to hold weight. We'll make more new content on a per-month basis than all of recorded human history. The old regime must be thrown away to accommodate the radically new world we're entering.
We'll land somewhere in-between, and I'm hoping it's much closer to (or even precisely) the latter.
I am not fighting for the smaller players but for large enterprises. That is illogical.
Well, this is the crux of the matter, isn't it? Do you, a human, have the right to look at copyrighted works and learn from them? Do you have the right to use AI to do the same?
- If you bought the book, you can read it.
- If the book is free, you can read it.
- If the painting is in a museum, or on Wikipedia, you can visit it.
- If Bozo the clown says you’re not allowed to look at drawings he posted online, it’s ok. You still can.
Same for AI.
That's what they did!
It was in fair use. So yes, they did have the right to legally train the data on copyrighted images.
Many artists don't believe this and the law is very much unclear.
In many cases the AI generated work literally looks like a clone.
Sure, but you still just cannot output anything that looks like a derived or copied work.
So, maybe ... how about if image generation nets hold onto the training images so that it can compare the generated output against its training data to ensure that it is not too similar.
/s (but only a little)
How much of that is going back to the copyright holders whose work their service derives value from ?
how much of the earnings of the student of art goes to the textbook authors, paintings and learning materials he used to get to where he is today?
But the AI as a whole is capable of reproducing the original in a recognizable form, and it does so on demand quite easily, because it was trained on them - how is it different than selling a zip file containing millions of copyrighted works, and also a bunch of new stuff?
so you're saying that the digits of pi is violating copyright then?