As a (creative) friend of mine flatly said, they refuse to use an LLM until it can prove where it learned something from/cite its original source. Artists and creatives can cite their inspirational sources, while LLMs cannot (because their developers don't care about credit, only output) by design. To them, that's the line in the sand, and I think that's a reasonable one given that not a single creative in my circles has been cut payment from these multi-billion-dollar AI companies for the unauthorized use of their works in training these models.
"See, those developers themselves have used CoPilot, so they approve the copyright infringement."
Even humans have a lot of internalized unconscious inspirational sources, but I get your point.
Regardless, deep learning models are valuable because they generalize within the training data to uncover patterns and features and relationships that are implicit, rather (simply) present with the data. While they can return things that happen to be within the training set, there is no reason to believe that any particular output is literally found there or is something that could be attributable, or that a human would ever attribute. Human artists also make meaning from the broad texture of their life experiences and general diffuse unattributable experience of culture.
Sure, this is something a random artist is unlikely to know, but if they are simply refusing to pick up a useful tools that can't give credit--say avoiding LLMs for brainstorming, or generative selection tools for visual editing, or whatever, their particular careers will be harmed by their incurious sentimentality, and other human artists will thrive because they know that tools are just tools, and it is the humans using the tools that make meaning that people care about.
[1] https://arxiv.org/abs/2504.07096
Why? Was it legal for me to download copyrighted songs from Limewire as "fair use"? Because a few people were made examples of.
I'm a musician, so 80% of the music I listen to is for learning so it's fair use, right? ;)
I would be happy with that outcome. I’m a fanfiction writer, and a lot of the stories I read are very much for learning. ;-)
[0] https://torrentfreak.com/meta-says-it-made-sure-not-to-seed-...
Secondly, there's an argument that the infringement happens only when the LLM produces output based in part of whole on the source material.
In other words, training a model is not infringing in itself. You could "research" with it. But selling the output as "from your model" is highly suspect. Your business is then based on selling something based other people's work, that you do not have rights to.
What fair use? Were the books promised to them by god or something?
True, but not the only relevant thing.
If the output of the LLM is "not very different from the original work" then the output could be the infringement. Putting a hypercomplex black box between the source work and the plagiarised output does not in itself make it "not infringing". The "LLM output as a service" business is then based on selling something based other people's work, that they do not have rights to.
It's falling for misdirection, "pay no attention to the LLM behind the curtain" to think otherwise.
I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original.
But I think that this is legally undecided and won't be decided by you or me, and it is going to be a more interesting and relevant question than "is the LLM model is very like the original work", which it clearly isn't. That's like asking "is this typewriter like this novel?" It can't be, but the words that came out of it could be.
Music has ended up in a place where short audio snippets are protected by copyright and must be licensed; but for short snippets of text the precedent has generally been that the copying needs to be more substantial. Distributed microplagarism of short phases might end up being ruled to be legal, even if wholesale reproduction is not. Which may not give copyright protection to the generated works, of course, as the question of machine authoring is entirely distinct.
That’s like saying the dictionary is micro-plagiarism of a huge number of sources because it uses all the words from those sources.
Plagiarism isn’t necessarily copyright infringement, and plagiarism isn’t illegal. Copyright infringement is.
Even still, your argument that everyone who generates 2,000 words in the style of (author) is plagiarizing is also flatly false. By that standard all English essays that mimic someone else’s style would be plagiarism.
The output of a LLM, when based heavily on that same page, pretends to be something novel. IMadeThis_Meme.gif
What am I, if not an LLM, ingesting copyrighted materials so that I may improve my own future outputs? Why is my own piracy not protected in the same manner?
You aren't a multi-billion dollar company
The Berne convention mentions "fair practice", and puts the responsibility on the individual countries.
Where’s the threshold for forcing AI companies to retrain models without specific copyrighted works in them?
We need to frame this case - and ongoing artist-vs-AI-stuff -using a pseudoscience headline I saw recently: 'average person reads 60k words/day'.
I won't bother sourcing this, because I don't think it's true, but it illustrates the key point: consumers spend X amount of time/day reading words.
> It seems like the authors are setting up for failure by making the case about whether the AI generation hinders the market for books. AI book writing is such a tiny segment what these models do that if needed Meta would simply introduce guard rails to prevent copying the style of an author and continue to ingest the books.
and from the article:
> When he turned to the authors’ legal team, led by high-profile attorney David Boies, Chhabria repeatedly asked whether the plaintiffs could actually substantiate accusations that Meta’s AI tools were likely to hurt their commercial prospects. “It seems like you’re asking me to speculate that the market for Sarah Silverman’s memoir will be affected,” he told Boies. “It’s not obvious to me that is the case.”
The market share an author (or any other artist type) is competing with for Meta is not 'what if an AI wrote celebrity memoirs?'. Meta isn't about to start a print publishing division.
Authors are competing with Meta for 'whose words did you read today?' Were they exclusively Meta's - Instagram comments, Whatsapp group chat messages, Llama-generated slop, whatever - or did an author capture any of that share?
The current framing is obviously ludicrous; it also does the developers of LLMs (the most interesting literary invention since....how long ago?) a huge disservice.
Unfortunately the other way of framing it (the one I'm saying is correct) is (probably) impossible to measure (unless you work for Meta, maybe?) and, also, almost equally ridiculous.
Legal cases are often based on BS, really an open form of extortion.
The plaintiffs might've been hoping for a settlement.
Meta could pay $xM+ to defend itself.
Maybe they thought Meta would be happy to pay them $yM to go away.
The reality is, there's very little Meta couldn't just find a freely available substitute for if it had to, it might just take a little more digging on their end.
The idea that any one individual or small group is so valuable that can hold back LLMs by themselves is ridiculous.
But you'll find no end to people vain enough to believe themselves that important.
To make fair use of a book's passage, you have to cite it. The except has to be reasonably small.
Without fair use, it would not be possible to write essays and book reviews that give quotes from books. That's what it's for. Not for having a machine read the whole book so it can regurgitate mashups of any part of it without attribution.
Making a parody is a kind of fair use, but parodies are original expression based on a certain structure of the work.
That's not true. That's what's required for something not to be plagiarism, not for something not to be copyright infringement.
Fair use is not at all the same as academic integrity, and while academic use is one of the fair use exceptions, it's only one. The most you would have to do with any of the other fair use exceptions is credit where you got the material (not cite individual passages), because you're not necessarily even using those passages verbatim.
- of a commercial nature;
- plagiarism;
- substantially large (e.g. whole work);
you're not on good legal footing.
Fair use and plagiarism are related, but they are two separate things. Especially when talking about the legalities of things, as we are here, it's vital to be clear and accurate about what specific legal issues are under discussion.
Facebook isn't claiming fair use to write academic papers about the books they're taking parts from. They're claiming fair use to feed them into LLM training. If that is a usage that is deemed to fall under fair use, then it won't require specific citation, even if it requires attribution in the more general sense (ie, crediting all the works you fed into your word-chipper), any more than you're required to cite specific passages when you're making a wholesale parody of a copyrighted work (also fair use) or writing a fanfic based on it (also fair use).
Anyway, the organizations behind popular LLMS are perpetrating massive copyright infringement and plagiarism. A fair use defense is rationally not possible.
Firstly, to claim fair use while you are making commercial use of the material is going to be difficult right off the bat. Fair use claims are bolstered by non-commercial use.
Small size of excerpts favors fair use claims. Can't do that here.
Transformative use? Not really; LLMs spit out the information verbatim if prompted in the right ways. Exploitative uses will probably not be viewed as transformative by the court.