Or viewing it on my television requires multiple copies as the data moves through wires and tuners and memory and storage. Obviously a video just be copied into RAM to process and display for normal viewing. Why would making a copy in memory to run through a training algorithm be different than an algorithm to encode or decode for display on a screen?
I don’t think transmission should count as making copies.
It’s different to store material vs stream and process.
“Copies” are material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device.
17 USC 101.There’s no way to look at your brain to perceive, reproduce, or communicate a work that you looked at, so under US law, your neurons are not a copy.
As for decoding a video to display it on a television, yes, that process requires making copies. That is why if you don’t have a license to a work, then decoding and displaying it is copyright infringement.
But I think that’s immaterial as I think the answer is that it doesn’t matter if you train on copyrighted work. Or maybe better that you don’t need a special license.
If I steal books and train on them, then I think that’s copyright infringement not because of the training but because I made an infringing copy. However, if I have a license already to read those books (ie, I bought a copy at a book store) then it’s not infringement to train an AI, or loan it to a million people or whatever I like with it as I bought a copy.
Yes but you’re not allowed to copy and distribute so you’re never going to actually do that. Also you’ll pass away with the knowledge you gained unlike the AI you’re training it on which will contain remnants of the work in all copies of itself, however abstracted away you consider it.
I can very legally loan my book or movie to one person a week for all perpetuity so it could be a million people. Does the actual number matter?
On the other side if you just dump tons of artist into a model to teach it how to draw, it'll be a much harder argument to make when the result doesn't really look like any artist in particular.
However a big problem with current models is that their vocabulary is still quite limited, so the only way to get some interesting art style with just a text prompt is to have a "by <artist name>" in the prompt. That's rather fishy. Meanwhile if you custom train a model, the amount of images you put into it will be drastically lower and more focused than what went into into initial training, so that's a bit fishy too.
Long term I expect art models to filter trademark terms and just get better language understanding or other means of input (e.g. sketches in ControlNet), so that the model ends up being unable to reproduce something close to an existing work out of the box. It wouldn't stop the user from using that model to perform some copyright violation, but it would require intend and effort on the users part.
The problem with the music industry is the industry part : it's been generating standardized, and easily mass produced music for the past 20 years. You can't expect the process to not be fully automated eventually ( i've already felt like some pop or rap song could have been for a long time).
You can't expect it not to backfire at some point. And IMHO this is well deserved.
And it’s going to be even worse with multimodal ai and robotics : imagine a robot walking in the streets, looking at cars, and ads, and people, then generate original content based on what it saw on the street (much like a human do). Who’s going to collect for all the licensed content it was exposed to ?
That's going to be totally impossible to apply. Imagine a sound processing software than "learned" mixing based on a collection of popular records, and provide you, the composer, with presets => wrong, copyright.
Ok, now imagine an ai that learned how to configure a sound processing software based on records. It only outputs numbers and presets for the software. => wrong, copyright.
Now, imagine an ai that learned those settings, by asking hundreds of human what they like in some kind of A/B testing. The humans will judge based on the songs they've heard in their lives, and the end-result (the settings) will basically be the same as in the previous example. But somehow, this will be fine.
Honestly, i don't see how this can fly.
Those AI produce original content. The means by which they produce it shouldn’t be relevant.
The means through which a piece of content has been produced is totally irrelevant. The original content has to be distinguishable in the new one for copyright laws to apply. This has been judged over and over when sampling became more and more widespread. Some DJ sampled the original songs blatantly and thus had to pay the original authors. Others applied so many sound transformations that it was basically impossible to identify and thus copyright was basically impossible to apply.
“If it could learn openai would create a new programming language, train it for real, and it would be the greatest programmer.” I don’t understand what you’re saying here. It can invent new languages, but they would those be better than existing ones? It’s less intelligent than a human (though still intelligent) so the language would be worse than existing ones
[1]: https://arxiv.org/abs/2012.07805 [2]: https://arxiv.org/abs/2301.13188
I used to think “well this dumb law can’t be so bad if it kills Facebook that’s a net positive” but I don’t think having directions that allow Beatles cover bands like Oasis but disallow an AI Oasis will be good for society. Comically, I think this because one day we’ll be digital beings and it will be funny if these old timey 21st century laws only allow creativity from meat-based consciousnesses because my digital consciousness is basically just an advanced AI that I legally treat as me.