> The firm said payment for all of the copyrighted material already used in LLMs would cost the companies that built them "tens or hundreds of billions of dollars a year in royalty payments."
https://www.businessinsider.com/marc-andreessen-horowitz-ai-...
Watching the superstars of venture capital whine that copyright is unfair is quite something, though.
It seems like the only people that will be impacted are the abusive auto-generated spammer accounts with thousands of garbage tracks uploaded garnering 1200 streams a year by people accidentally playing them via Google Home misinterpretations etc.
So yes an individual song might be $3 per year but that just shows how poor their royalties are to begin with. And tries to obscure the fact that artists don’t just release one song ever.
There’s thousands of artists who maybe even were somewhat successful at some point in their career but would have a lot of songs in their back catalog that don’t get that many streams annually. Suddenly they’ve gone from not making enough per stream from Spotify, to just getting paid nothing at all.
The fact that "research, technological advancement, and construction would proceed much more quickly" without copyright is exactly why abolishing copyright is the just and correct thing to do.
I am not sure where the “unnatural restriction” quote comes from, could you illuminate me?
As far as what the constitution says about copyright, it seems only to say that it is “To promote the Progress of Science and useful Arts”
which, as an involuntary donor, is exactly what I want
- slavers, probably.
Of course slavery != AI, but the argument that we should protect companies from their expenses to enable their bad business model is very entitled and presumptuous.
Thousands of companies have failed because their businesses models didn’t work, and thousands more will.
AI will be fine. It probably won’t be as stupidly lucrative as the current model, but we’ll find a way.
If only saying it would make it so.
Unfortunately, it's not easy to make this legal argument given how copyright law only protects fixed, tangible expressions, not ideas, concepts, principles, etc. and has a gaping hole called 'fair use.'
1. Models are trained with 100% uncopyrighted or properly licensed input data
2. Every output of the ML model is evaluated to make sure it's not too close to training data
3. Copyright law is changed to have a specific cutout for AI
#1 is the approach taken by Adobe, although it generally is harder or more expensive to do.
#2 destroys most AI business models
#3 has been done in some countries, but seems likely that if done in the US it would still have some limits.
For example, I could train a model on a single image, song, or piece of written text/code. Then I run inference, and get out an exact copy of that image, song, or text. If there are no limits around AI and copyright, then we've got a loophole around all of copyright law. I don't think that the US would be up for devaluing intellectual property like that.
4. A ruling comes down that enshrines what all the big companies have been doing (with the blessings of their armies of expensive, talented, and conservative legal teams) as legitimate fair use
https://twitter.com/Rahll/status/1739003201221718466
It's frankly impressive how well this image is embedded in the weights of their model, down to tufts of hair. And it's far from the only one.
You are forgetting the massive asterisk that you need to provide multiple paragraphs of the original article in order to get verbatim output from chatgpt. In what world are people doing that to avoid paying the NYT?
https://arstechnica.com/tech-policy/2023/12/ny-times-sues-op...
And this is a screenshot of their session whith copilot
https://cdn.arstechnica.net/wp-content/uploads/2023/12/Scree...
Its funny that the behavior was patched if OpenAI believes it isn't copyright infringement.
It supports an argument that GPT shouldnt produce outputs that are extremely similar, not that the content can not be used as an input.
as well as all the output it ever generated
Does that mean that models that can not produce copies of X length ARE fair use?
not necessarily
"sufficient but not necessary" I believe is the term
it is a fringe case that rarely occurs, and only with a lot of user prompting.
legal discovery could almost certainly compel the LLM host to provide access to the output of the weights themselves without the "gatekeeper" present
If there is a case to be made, I think it has to be around the original use of the works, the transcription process. Not the weights, or the output
If the archive can't produce the original work, it's not infringing.
If you printed the binary of Harry Potter and sold it as a painting, that would be fair use. It doesn't matter if the data is encoded in it, if it is not used for extraction.
Think of Andy warhol's Campbell Soup. Nobody is going to confuse the art for a can of soup and try to eat it. That's not being sold as a label for other soups. However, the original Campbell Soup data is absolutely encoded there.
It's actually just "anyone making models". If you train a model with other people's art (without their permission) and then distribute the model or output for free, your still stealing their work, even if you make zero profit.
Yes, I know Adobe said so. No, I don't trust them.
Facts:
1. Adobe Firefly is trained with Adobe Stock assets. [1]
2. Anyone can submit to Adobe Stock.
3. Adobe Stock already has AI-generated assets that are not correctly tagged so. [2]
4. It's hard to remove an image from a trained model.
Unless Adobe carefully scrutinize every image in the training set, the logical conclusion is Adobe Firefly already contains at least second-handed unauthorized images (e.g. those generated by Stable Diffusion). It's just "not Adobe's fault".
[1] https://www.adobe.com/products/firefly.html : "The current Firefly generative AI model is trained on a dataset of licensed content, such as Adobe Stock, and public domain content where copyright has expired."
[2] Famous example: https://twitter.com/destiny_thememe/status/17448423657672255...
However, if other artists are inspired by this style of comic, and it influences their work - that is simply fair use. If that artist is some rando using a tool like Midjourney - that is inspired by the art but doesn't reproduce it - it is not at all clear to me that this is not also fair use.
That already clearly means that they couldn't publish the model directly even if they wanted to, since they don't have the right to distribute copies of those works, even if they are represented in a weird lossy encoding. Whether it's legal for them to give access to the model through an API that prevents returning copyrighted content is a much more complex legal topic.
or… it is possibly a transformative work of the all the works that went into it's training, which would lead to a strong fair use argument. Given how permissive the courts have been with transformative work, this seems like an easier argument to make.
Human artists can also create copies when instructed to.
That doesn't mean that the rest of their work isn't transformative, or that the process of leading isn't fair use.
Similarly, the law doesn't bar artists from learning, but provides recourse, when and if artists create and sell copes.
Of course. The model isn't making a decision as to what may be used as training data. The humans training it do.
>If someone uses Midjourney to produce someone else's IP, that user (not Midjourney) would be in violation of copyright
That's like saying that if a user unpacks the dune _full_movie.zip I'm sharing online, it's them who have produced the copyrighted work. And me, the human who put the movie Dune into that zip file, is doing no wrong. Clearly, there is no compression algorithm that can launder IP, right?
>However, if other artists are inspired by this style of comic, and it influences their work - that is simply fair use
The AI isn't inspired by anything. It's not a sentient being, it's not making decisions, and its behavior isn't regulated by laws because it does not have a behavior of its own. Humans decide what goes into an AI model, and what goes out. And humans who train AI models on art don't get "inspired". They transform it into a derivative work — the AI model.
One that has been shown to be awfully close to dune_full_movie.zip if you use the right unpacking tools. But even that isn't necessary. Using work of others in your own work without permission and credit usually goes by less inspiring words: plagiarism, theft, ripping off.
Regardless of whether you reproduce the work 1:1, and whether you can be punished by law for it.
>tool like Midjourney - that is inspired by the art but doesn't reproduce it
Never in the history of humanity has the word inspired meant something that a tool can do. If it's "inspired" (which is something only sentient beings can do), then we should be crying out about human right abuses the way the AI models are trained and treated.
If it's just a tool, it's not "inspired".
You can't have your cake and eat it too. Either pay your computer minimum wage for working for you, or stop saying that it can get "inspired" by art (whether it's training an AI model or creating a zip file).
Like leaded gas, the government can make regulations to deal with anything should they choose to.
Why do you think it's okay for massive companies to freely profit off the work of others?