Training AI Using 'Pirated' Content Can Be Fair Use, Law Professors Argue
torrentfreak.com
torrentfreak.com
You are allowed to read books from the library and remember what they say, and use that information to inform your own future writing or speaking or actions.
You are allowed to listen to copyrighted music and learn from it. You can even play songs from The Beatles or Metallica in your garage as training.
You absolutely do have the right to "train your brain" on copyrighted material. Copyright restricts who is allowed to publish the work, not who is allowed to consume it.
You aren't allowed to download a torrent of pirated books as these companies have done and freely distribute it to multiple brains to train on.
If the brains can then write down the original works from memory, you aren't allowed to make copies of these brains and freely distribute them either.
You could count the words in a book and publish the word count, and while the information is based on the contents of the book, that would fall incredibly short of being a derivative work.
I suspect they committed whatever copyright violation is committed when they downloaded the copyrighted works. Training an AI on them is simply not related to the protections that copyright offers.
The syntax brought to mind Stephen Colbert's 2007 book, I Am America (And So Can You!)
The exact lines are a bit blurry and subjective, but this is a decent overview: https://www.lib.uchicago.edu/copyrightinfo/fairuse.html
If you're charging people to watch in a movie theater, you can't pretend that's educational. But if you want to show a video in a classroom that's relevant to the coursework, it's absolutely covered under fair use.
These guys are part of the problem IMHO. Having studied some law at university and then reading all of Lawrence Lessig's works when he dissected IP law 20 years ago for the Creative Commons project, I was left with the distinct feeling that "IP law" is ugly, unfair, arbitrary, ineffective, crippling to the mind, and devastating to progress and the economy.
I now strongly agree with the likes of Richard Stallman that "intellectual property" is a grotesque and bankrupt mess that should be avoided if you want to have any semblance of a mature conversation.
Actually, people training AI have a point. New technologies expose the silliness of bad ideas we've clung to for three centuries. And the only way forward is to hugely reform or repeal most of IP law for EVERYONE.
The truth is that long before the Statute of Anne "copyright" has its roots in censorship and political control. We used to burn the printers of "seditious works" on pyres of their own books in St James' square in London.
Much of "IP law" still functions the same today but hidden behind a cover story about "protecting creators".
I now mentally substitute the phrases "intellectual property" with "coercive control of information" (CCI).
CCI gets to the nub of power relations instead of pretending we have nice IP "laws" that are applied uniformly. In reality copyright, patents and trademarks have become tools of censorship and denial for those with money, and they do almost nothing to protect individual creators. Things like the DMCA are simply monstrous. If we can't reform and enforce it to actually protect creators it's time to scrap the whole rotten show in my opinion.
Nobody is prevented by trademark law from offering any product or service, including commercially; the trademark holder’s consent or an applicable exception is only needed in order to use the trademarked name/logo in their own offering’s name/logo.
(agreeing with you here and suggesting that people who work with consumer protection marks should distance themselves from the term "IP" entirely)
Only the very wealthiest would be allowed to produce with it, they would form state backed cartels and we would be no better off for its invention.
Fair use is an affirmative, after-the-fact defense.
Using copyrighted material in a fair use way seems fine to me and is important. But these companies not wanting to pay for creating their model and then just claiming fair use is silly.
Certainly model developers would prefer not to pay if given the option, but I also feel it's not untruthful to say that it hasn't actually been feasible to license content on the scale required.
Even just for training an object detection network as a side project, I struggled to find sufficient pre-training material outside of web-scraped datasets like ImageNet. I even contacted Getty and was told directly that they don't license images for machine learning.
Something like a compulsory licensing scheme where you pay into a pot to train a model could potentially work. Mostly, I hope whatever we eventually get is feasible for open source groups, individual developers, universities, smaller companies, etc. rather than only being made with the few biggest companies in mind.
"simply don’t train anything" does not seem ethically (many models have the potential to or already are improving lives), politically (staying in the lead is currently seen as an important issue), or legally (as noted in this brief, "the ultimate test of fair use is whether the copyright law’s goal of ‘promoting the Progress of Science and useful Arts’ ‘would be better served by allowing the use than by preventing it.’") viable to me.
No conflict of interest here ? Asking for a friend. /s