What would that email look like? "Here's some lyrics you never wrote and some audio you've never heard. Can I use it in my song?"
https://genius.com/Kanye-west-kanye-west-and-taylor-swifts-f...
(For those that would not know, he coined the term "robot" from the slavic root "work".)
For somebody like David Guetta, it always has been financially possible.
they're flouting a sample clip with eminem on their homepage.
this just seems like clever marketing to me.
Does Supabase owe him royalties when they use that same model that trained on Eminem’s lyrics to make a chat bot that can interact with their support documentation?
Its entirely fair game for subjects in a given field to protest when an AI is trained on their work and then is used to power competition to them without any form of remuneration or licensing to compensate them.
Logically, the model must be considered in fair use if we are to allow Supabase Clippy. This is how the courts will see it!
We have yet to see this play out in court, but we likely will in the following years.
It took the Supreme Court to overrule, where they created the legal doctrine of “commercially significant non-infringing use” as part of determining fair use. They looked at the VHS rental market and the numerous other industries that were dependent on this technology. The courts do not want to be in a position where a ruling can have a profound impact on commerce especially when balanced with extending copyright claims at the detriment to freedoms of speech.
This is why legal experts keep saying the model and the outputs will be considered different and why the model itself is fair use and why the risk-averse lawyers at Microsoft are getting behind these tools… and even why they are letting people build on top of them before they were ready for prime time, to flood the market with non-infringing uses!
In their capacity as tools, LLMs probably would fall under a similar model. However, the LLM itself is substantially based on the IP of all of the creators of the data used in its training set. This may or may not be granted fair use status, or it may or may not even be seen as infringing on the copyright of those works at all - but either way, it is not covered by that case as a legal question (at least in my own non-lawyer interpretation).
Or I guess you can give me a little bit of time to go and help you do some of your own research, which I will do right now, so just hold on a bit!
I'm not trying to trick people or win some hypothetical argument, I'm trying to help people see how the courts will consider these issues and how and why we should agree with their rulings that these tools are fair use of copyright protected works!
I love copyright, probably more than most people on these forums, but the liability needs to be on the people using these tools. Supabase Clippy is fantastic and they should not bear any costs, even implicitly through OpenAI paying royalties. Someone who releases a "Sarah Anderson Cartoon Maker" tool using Stable Diffusion should still be found to wrong Sarah Anderson's common law right to publicity just as someone who releases a "Vacation Photo Background Cleaner-Upper" tool using Stable Diffusion should not bear any costs, even implicitly through StabilityAI paying royalties to Sarah Anderson.
These must be considered in their capacities as tools regardless of how they were made and for reasons foundational to the law itself: How can you prove that I used a given tool such as Stable Diffusion or the "Vacation Photo Background Cleaner-Upper" tool without concrete evidence of the tool, such as the presence of the software on my laptop? Can law enforcement get a warrant to search based on literally no visual evidence that Sarah Anderson's works were somehow used in the training process of a tool that I have on my private property?
Edit:
---
In the Texas Law Review in March, 2021, Mark Lemley, a Stanford law professor, and Bryan Casey, then a lecturer in law at Stanford, posed a question: "Will copyright law allow robots to learn?" They argue that, at least in the United States, it should.
"[Machine learning] systems should generally be able to use databases for training, whether or not the contents of that database are copyrighted," they wrote, adding that copyright law isn't the right tool to regulate abuses.
But when it comes to the output of these models – the code suggestions automatically made by the likes of Copilot – the potential for the copyright claim proposed by Butterick looks stronger.
"I actually think there's a decent chance there is a good copyright claim," said Tyler Ochoa, a professor in the law department at Santa Clara University in California, in a phone interview with The Register.
In terms of the ingestion of publicly accessible code, Ochoa said, there may be software license violations but that's probably protected by fair use. While there hasn't been a lot of litigation about that, a number of scholars have taken that position and he said he's inclined to agree.
https://www.theregister.com/2022/10/19/github_copilot_copyri...
https://texaslawreview.org/fair-learning/
Both Lemley and Ochoa state that the models themselves are probably protected by fair use. Meaning, it is perfectly fine for OpenAI to train their models on publicly accessible copyright protected works without asking for permission or having to pay any royalties or having to adhere to any of the terms of the license.
They are also free to distribute this tool and to charge for people to use this tool.
What Ochoa is saying about there being a good chance of there being a copyright claim is that this tool doesn't absolve the users of the tool from copyright violation. The liability is on the person using the tool regardless of the tool being used. It's not the intent that matters with copyright, it's that you ended up publishing something that looks enough like someone else's picture that twelve people consider it to be not too different than a simple photocopy.
Now, this can still be a problem for Copilot because an engineer's company might not want to be injecting a lot of copyright protected code into their products, but for the most part the outputs from Copilot have been non-infringing. That it sometimes produces infringing code does not matter to anyone other than the person using Copilot.
Ochoa goes on in detail about what is and isn't covered by copyright with regards to code, which is one of the things that gives me confidence to use Copilot and know that I'm not putting myself at risk:
But in terms of where Copilot may be vulnerable to a copyright claim, Ochoa believes LLMs that output source code – more so than models that generate images – are likely to echo training data. That may be problematic for GitHub.
"When you're trying to output code, source code, I think you have a very high likelihood that the code that you output is going to look like one or more of the inputs, because the whole point of code is to achieve something functional," he said. "Once something works well, lots of other people are going to repeat it."
Ochoa argues the output is likely to be the same as the training data for one of two reasons: "One is there's only one good way to do it. And the other is [you're] copying basically an open source solution.
"If there's only one good way to do it, OK, then that's probably not eligible for copyright. But chances are that there's just a lot of code in [the training data] that has used the same open source solution, and that the output is going to look very similar to that. And that's just copying."
In other words, the model may suggest code to solve a problem for which there's only really one practical solution, or it's copying from someone's open source that does the same thing. In either case, that's probably because a lot of people have used the same code, and that shows up a lot in the training data, leading to the assistant regurgitating it.
So in practice it is pretty easy to tell that Copilot is spitting out purely functional suggestions basically all of the time as there isn't really any other way to wire up a unit test or call a specific API.
Ironically, if Copilot gets better at "software architecture" then it starts to cross over into the expressive parts of software that are indeed covered by copyright, meaning these issues of liability become harder to discern to the end user and enough of a problem that GitHub would want to figure out attribution or somehow "clear" the suggestions for the user.
Where we differ is in how certain we are that the IF is true. I for one believe there is a good chance that the training of an LLM on copyrighted works does infringe on the copyright of those works (if no other exceptions apply, such as the LLM being trained only for academic research purposes, of course).
My original response:
Where Sony v Universal definitely applies though is when evaluating whether OpenAI's selling of GPT3 to others who then use it to create copyright-infringing works would make OpenAI liable for contributory infringement. Here, the similarities are crystal clear, and the conclusion is simple: since there clearly exist non-infringing uses of GPT3 (such as Supabase Clippy), OpenAI is fully in the clear to sell GPT3, just as much as Sony was for selling the VCR.
However, this assumes that OpenAI has the rights to the IP of GPT3 itself in the first place, which is a prerequisite to them being allowed to sell it at all. Sony certainly had the rights to the IP of the VCR - Universal never claimed that the VCR was a derivative work of their movies.
Essentially, in Sony v Universal, Universal was claiming (1) that Sony was liable for contributory infringement, since (2) all customers of Sony who used it to record and then playback a Universal show were guilty of copyright infringement. The court established that (2) was in fact fair use, and from there automatically (1) become false, since now there was an established legal way of using the Sony product.
But, in a hypothetical OpenAI v Universal, Universal could plausibly claim that (1) OpenAI is liable for copyright infringement directly, since they are distributing GPT3 , (2) which is a derived work of Universal's IP used in the training set of GPT3.
I'm more certain because I'm thinking about this separately from my opinions as to how the courts will rule. Yes, I will just so happen to agree with that ruling because I happen to agree with the logic and knowledge contained in our legal process. I agree that the existing legal doctrines already capture the spirit of what we are asking them to judge. I agree that the statutory law, case law and doctrine that informs their judgement will successfully balance both the limited rights of copyright holders and the natural rights of a public to unburdened access to the arts, knowledge and information. I agree with their process of balancing the potential impact on existing commercial practice with the potential impact on new forms of commercially significant non-infringing practices.
Some things in copyright might just seem unfair, like the case of Baker v Selden:
In 1859, Charles Selden obtained copyright in a book he wrote called Selden's Condensed Ledger, or Book-keeping Simplified. In it the book described an improved system of book-keeping. The books contained about twenty pages of primarily book-keeping forms and only about 650 words. In addition, the books contained examples and an introduction. In the following years Selden made several other books, improving on the initial system. In total, Selden wrote six books, though, evidence suggests that they were really six editions of the same book.
Selden, however, was unsuccessful in selling his books. He originally believed he could sell his system to several counties and the United States Department of the Treasury. Those sales never happened. Selden was forced to assign his interest—an interest that apparently was returned to his wife after his death in 1871.
In 1867, W.C.M. Baker produced a book describing a very similar system. Unlike Selden, Baker was more successful at selling his book–selling it to some 40 counties within five years.
Selden's widow, Elizabeth Selden, hired an attorney, Samuel S. Fisher, a former Commissioner of Patents. In 1872, Fisher filed suit against Baker for copyright infringement.
The poor old widow lost. Boohoo. But this was a just ruling!
I'm not sure that the people who think that ChatGPT is guilty of copyright infringement are thinking about the issue in a balanced manner. Luckily our courts probably will!
One strategy that the defense could use to lower their risk profile is to allow open access to their models and allow an entire ecosystem of commercially significant non-infringing uses to blossom because they are aware of how the courts will be influenced based on existing statutory and legal doctrine...
The training data is also publicly available, https://pile.eleuther.ai/
https://arxiv.org/pdf/2101.00027.pdf
7.1 Legality of Content
While the machine learning community has begun to discuss the issue of the legality of training models on copyright data, there is little acknowledgment of the fact that the processing and distribution of data owned by others may also be a violation of copyright law. As a step in that direction, we discuss the reasons we believe that our use of copyright data is in compliance with US copyright law.
Under pre (1984) (and affirmed in subsequent rulings such as aff (2013); Google (2015)), non-commercial, not-for-profit use of copyright media is preemptively fair use. Additionally, our use is transformative, in the sense that the original form of the data is ineffective for our purposes and our form of the data is ineffective for the purposes of the original documents. Although we use the full text of copyright works, this is not necessarily disqualifying when the full work is necessary (ful, 2003). In our case, the long-term dependencies in natural language require that the full text be used in order to produce the best results (Dai et al., 2019; Rae et al., 2019; Henighan et al., 2020; Liu et al., 2018).
Copyright law varies by country, and there may be additional restrictions on some of these works in particular jurisdictions. To enable easier compliance with local laws, the Pile reproduction code is available and can be used to exclude certain components of the Pile which are inappropriate for the user. Unfortunately, we do not have the metadata necessary to determine exactly which texts are copyrighted, and so this can only be undertaken at the component level. Thus, this should be be taken to be a heuristic rather than a precise determination.
1984. Sony corp. of america v. universal city studios, inc. 2003. Kelly v. arriba soft corp. 2013. Righthaven llc v. hoehn.
Some countries will try to hamstring AI development by saddling it with perversions of copyright use, and those countries will increasingly fall behind technologically and economically.
According to a quick google search, even cover bands performing live are required to pay royalties to the original artists whose songs they're performing.
Maybe there's some loophole when using an artist's likeness while changing the actual music and lyrics, I don't know.
But "protected speech" isn't meant to be a way to make money off of other people's work.
Curious what part of "Midler v. Ford Motor Co" you think wouldn't apply to a musical performance.
It's more than that. If you were part of a band and now you are solo (plenty of cases), you need to pay your old band if you play an old song you were part of (unless you own 100% of the lyrics and music rights, which is rare).
By the way, I am totalyl OK with this deal/agreement/way of doing things.
Do I need to ask permission to impersonate Eminem during a comedy performance? How about during a musical performance?
The answer is no. This is fair use 101.
If you cover a song of his while doing an impersonation then this is covered by the performance fees the venue you’re playing at pays every year to ASCAP and BMI.
Interesting. I guess you consider the use of Eminem’s voice to be an ‘in the style of’. I think, being trained on Eminem songs, it’s more like a sample.
The example I like (even though Weird Al always obtains permission so doesn't need Free Speech protection) is "Smells Like Nirvana" compared to "Fat".
Smells Like Nirvana is satire. The thing criticised is Nirvana's "Smells Like Teen Spirit", and the wider Grunge phenomenon it exemplifies. Al's song talks about how Kurt's lyrics are mumbled and hard to understand, "It's hard to bargle nawdle zouss / With all these marbles in my mouth", and about how it's too loud so it'll annoy your mom and dad, "We're so loud and incoherent / Boy, this oughta bug your parents". This can't work if Al uses Mariah Carey's "All I Want For Christmas" instead of Teen Spirit.
"Fat" isn't satire. Michael Jackson wasn't obese or even overweight, and his song isn't causing people to eat too much junk food, Al's song is just using the outline structure of Jackson's "Bad" because that fits and his new song is funny. If his lyrics had fitted better to Rick Astley's "Never Gonna Give You Up" that would have been fine.
The answer is maybe. See the case I cited above.
> This is fair use 101.
Voices are not copyrightable so fair use doesn't seem terribly relevant here.
This means he wasn’t paid for the performance, is that correct?
Because otherwise doing it for paid performances and promoting it on social media but saying “it’s not commercial bro” doesn’t seem very genuine
Midler pursued a common law judgment against Ford for using her distinctive voice without her authorization. The appellate court pondered the question of whether or not an artist's voice is a distinctive personal feature over which a person has controlling rights from appropriation. Midler was not seeking damages for copyright infringement of the song itself, but rather for the use of her voice which she claimed was distinctive of her person as a singer. The recognition of Midler's voice in the commercial was found to be the intentional motivation and a major feature of the commercial.
This is a common law right to publicity. This is our common law right to be in control of our likeness.
David Guetta is being paid but he is not being paid by the promoters or the audience because he says he is Eminem so there is no reasonable claim to a violation of Eminem's common law right to publicity.
Whereas with Midler, the commercial production crew was seen by the jury to be paying an artist to convince an audience that Bette Midler was endorsing the product.
These are completely different situations.