9k authors say AI firms exploited books to train chatbots
latimes.com
latimes.com
Fundamentally their asking to renegotiate the sale price of their work depending on what the purchaser decides to do with the book.
I'm looking forward to the time when I must declare if I'm purchasing a book for:
- personal enjoyment
- to produce a parody
- burn itYou could devise a scenario whereby an artist deliberately used pirated works in order to make a satirical[0] point about copyright itself (or the global publisher oligopoly, etc.) that might be accepted by a judge as a fair use argument, but that is not what's happening here.
[0] Parody is when you change a work to make fun of it (eg. a twilight parody where Edward and Bella's dialog is in Valspeak). Satire is when you change a work to make fun of something *else* (eg. a version of Twilight where Edward and Bella are replaced with Putin and Trump). The general reason Satire isn't as protected by Fair Use as Parody is because you can usually replace the (ab)used work with something else while still making a similar point about the actual subject of your critique (eg. make a Putin/Trump version of Pride and Prejudice instead of using Twilight).
Or am I way behind the curve on this one.
Here’s the four factor test courts use:
> In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include:
> 1. the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
> 2. the nature of the copyrighted work;
> 3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
> 4. the effect of the use upon the potential market for or value of the copyrighted work.
I think points 1 and 3 are the least debatable: commercial LLMs should easily fail 1 and all LLMs should pass 3. EDIT: actually changed my mind on this. If they’re using the whole work, then LLMs should actually fail #3. It’s not just how much of the model it makes which is insignificant relative to all the rest of the text in the model, but also how much of each individual work is used. If that’s 100% then that’s actually an easy fail.
We could have a debate about #2 till the end of time and I’m not really here for that, not today anyway, but it is probably worth debating.
#4 is also absolutely debatable, but I think there is a large potential for LLMs to harm the market for the original source material that they were trained on.
Three is confusing to me. When somebody make a Mona Lisa with a mustache, they've used the entirety. It's how much they modified that's important? Whether it can be mistaken for the original.
Nobody is going to mistake an AI-ingested corpus for the original. It seems quite debatable.