Training algorithms on copyrighted data not illegal: US Supreme Court
towardsdatascience.com
towardsdatascience.com
However, the 2nd Circuit's ruling is not binding on any other federal circuits.
Also, as Enginerrd stated, the holding is not nearly as broad as the article makes it out to be.
The holding was:
1. Google’s unauthorized digitizing of copyright-protected works, creation of a search functionality, and display of snippets from those works are non-infringing fair uses. The purpose of the copying is highly transformative, the public display of text is limited, and the revelations do not provide a significant market substitute for the protected aspects of the originals. Google’s commercial nature and profit motivation do not justify denial of fair use.
2. Google’s provision of digitized copies to the libraries that supplied the books, on the understanding that the libraries will use the copies in a manner consistent with the copyright law, also does not constitute infringement.
Based on the above holding, I think the article's conclusion is a stretch for general training algorithms using copyrighted data because: (1) there would not be a library supplying the information to the training algorithm, (2) there would be no similar display of snippets, and (3) we do not know if a training algorithm would provide a market substitute for the copyrighted data.
The HN link text:
> Training algorithms on copyrighted data not illegal: US Supreme Court
The sub-heading for the article:
> Training algorithms on copyrighted data is not illegal, according to the United States Supreme Court.
I wonder, does this mean I can scrape Instagram/Facebook for photos and use them for face recognition? Is that 'fair use'? Is an Instagram post a publication?
As I understand, it's generally viewed to be th case that a circuit split makes cert. more likely, sure.
> So while not legally binding, it might be in practice indicative.
I guess that it's indicative that, barring change in membership of th court, cert. would likely be denied in a future case raising the same issue from the same or a different circuit with the same result.
It definitely should not be seen as indicative of anything on the merits other than that the members of the court don't see it as obviously and urgently wrong.
Other than copyright there are privacy considerations too. For example Gmail's Smart Compose is trained on users' private messages so you don't want it to memorize "private" details (such as credit card numbers): https://arxiv.org/abs/1802.08232
Is it possible to solve this by adversarially checking if the output is "original" enough or not? Or is that intractable, given how much resource our society already pours into making the same classification in court?
Of course if you don't need a complete longitudinal patient chart then generating realistic data for just one aspect can be a lot simpler.
There is a single startup we're working with that has come much closer, but still isn't quite there. I think that if they manage to get funding they'll be somewhere usable in the near future.
If your model is any more useful than something trained on coarse aggregates, then it can be used to reidentify individuals. This is a pretty hard dilemma in the entire industry, not just health.
I hope my observations are skewed, but instead of trying to seriously address this issue I've seen an entire legal-loophole style data laundering industry emerge where highly identifiable information changes hands without it 'technically' changing hands in the legal sense. I'm talking about entities like datarepublic.
Is there a way to determine how much information about a particular individual has been leaked out?
That said, these many-pointed court tests make it really hard to do anything around copyrighted stuff without being a large corporation, which is problematic.
Good thing that these court tests are establishing that the argument, rather than being technical, is about creation of value to society. And about potential economic damage to the copyright owners. As long as there's no damage and value to society, it seems that the courts are fine with the use of data.
1) Federated machine learning. Basically you let google train their model on your private data, with the understanding that some of it would be used for the public goods. But this separate your data from the weights.
2) Train your own models, and ask google to open Gmail such that they will call your models.
They used some pretty reasonable tests of copyright infringement to conclude that no such infringement occured.
I think the logic makes sense because imagine if humans were prevented from getting ideas from watching movies. It seems similar to not letting AI watch every movie ever and learn.
Hence, they had to use fair use to justify it.
I think if you could train the AI without having to make a copy first, such as having the AI read the physical books directly, or in the case of your movie example having the AI watch the movie on a TV hooked up to a DVD player playing a copy of the movie on DVD that you bought from a retailer authorized by the copyright owner to sell such DVDs, you might not even need to make a fair use argument.
The definitions section of the US copyright statutes, 17 USC 101 [1], defines "copies" like this:
> “Copies” are material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term “copies” includes the material object, other than a phonorecord, in which the work is first fixed.
and "fixed" is defined like this:
> A work is “fixed” in a tangible medium of expression when its embodiment in a copy or phonorecord, by or under the authority of the author, is sufficiently permanent or stable to permit it to be perceived, reproduced, or otherwise communicated for a period of more than transitory duration.
An AI reading or watching the work as one of many many works in order to learn weights for a neural net does not result in a material object from which the work can be perceived, reproduced, or otherwise communicated. Thus, there is no copy, and hence no copyright issue.
This does not hold true in all cases. Note that the ruling lists the end goals as 'fair use' goals and that that seems to have been an important part in the conclusion reached.
The key thing to strive for in creating derivative works that are deserving of copyright protection in their own right is that they contain 'substantial originality', mere machine transformation does not qualify.
That minimises the contribution of thousands of researchers in designing the models and their training regimens.
Presumably this would require a digital camera capturing a copy of the book and storing it for some amount of time in computer memory, which seems equivalent to copying a digital version of a book or movie into the AI's computer memory.
It seems to me that "copy" is a legal term of art. Like, if I view a website, it might say I'm not allowed to copy the information. But depending on the level of abstraction, the data has been copied many times by many entities just to get to me, all the layers of machines, caches, retransmissions, etc. and exactly what I do with the normal functions of my browser cause more copies to be made.
Either this is not considered copying or it is considered fair use, but it seems pretty arbitrary to me, except that obviously considering it infringement would not advance the constitutional purpose of IP protections.
Is it legal to use the data to train an algorithm if the license disallows that explicitly?
In other words: you could not have come up with the model without the original data.
Receiving, strictly, no. However with digital material, most use involves copying; for legitimate copies that copying is covered by an implied license for the normal use of the work, for copies which are not themselves authorized, there is likewise no implied license for use.
Also, receipt of digital copies itself often involves copying directed by the receiver, which is prohibited, and may even involve a request from the receiver to the originator to make the copy under circumstances which would be viewed as knowing that it was unauthorized, which, may often, as a solicitation of an unlawful act, itself be illegal.
No, because then contractual ToS, not naked copyright law, will be at issue. Even if Google doesn't have the right terms to forestall this now, it's a trivial change for them to adopt.