If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.
And as long as OpenAI have an office in Japan they can absolutely legally train the models, no?
Also, the notion of the downloading itself being an illegal act is not universal as others have pointed out.
Making and having your own copies, and doing what you want with them, has always been fine. At worst it's a grey area, but in many cases it's been protected as fair use.
Whether people have been sued for downloading works they don't have the right to copy onto their machines is irrelevant to whether it is actually illegal. And it certainly has nothing to do with fair use, which is about copyrighted works that you actually do have some right to.
IANAL so I'm not going to tell anyone what does and what does not constitute fair use in what jurisdictions.
BTW people are being sued for distribution because they make great examples because their offenses and thus the damages are much greater.
I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption.
I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and the treatment of individuals in this case.
Unless people are typically punished for downloading and using copyrighted content, there is no such inconsistency. They are not, so there is not.
Copyright troll lawsuits have been fairly public and widely covered in the tech press, and most criminal prosecutions come with a formulaic, gloating press release from the law enforcement folks responsible. So it's pretty easy to follow this stuff.
The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing kids.
Everyone who was sued by the RIAA was done so for possessing the music in a publicly accessible method. Distribution was never actually proven in any of the cases (including the high profile losses). Defendants who argued that accessibility does not qualify as distribution actually won their cases. Most who argued against the validity of the evidence acquisition also won their cases.
'Acquiring' is more difficult to pursue legally. It's easier to go after distribution. In this case, Meta or OpenAI did not distribute anything because they are not chumps. They can go after whoever posted the dataset containing books. Not sure if that is eleuther or just some random person on the internet. In either case, the strategy of going after the rich companies won't work.
No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide.
A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.
And anyone using or training a model trained on works for which they do not have the rights? Completely fucked. Or at least, they must accept this as a real risk.
I won't be particularly thrilled if that turns out to be the case, but I wouldn't be surprised if it does.
But as you say, we won't know until it's tested in court. And even then, often court cases around a complex topic like this will end up with a ruling that only clarifies a narrow aspect of it. So it might take many related court cases before we have a pretty good understanding of where the law stands. And then, of course, the law could change.
Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.
What is the blackbox “limit” here? Is the mean value of all images in imagenet (which contains many copyrighted images) violating copyright? Is the character count of sarah silverman’s books? What about a prime number representing them - https://en.wikipedia.org/wiki/Illegal_number?wprov=sfti1
Training is much more similar to a character count than an illegal prime in my view, and thus, is almost certainly going to be okay/found to be okay. If not, something like, 90% of all models used today had some component trained on copyrighted data of some form.
Its unmistakably not a derivative work of the inputs individually or collectively, since a derivative work must be itself an distinct work of authorship (the same as the work of authorship requirement for copyright), and the output of a purely mechanical process is not.
The collection of inputs itself might be a derivative work of the individual inputs, before considering Fair Use.
Please omit flamebait and swipes, as the site guidelines ask: https://news.ycombinator.com/newsguidelines.html. Your comment would have been fine without that bit.
Ask it to write a book similar to Harry Potter, and it'll make an attempt at it. But human writers absolutely do read Harry Potter and write similar books, and that's perfectly legal. There have probably been thousands of published books inspired by Lord of the Rings.
So if it is satire, or uses an insignificant piece of the work within a larger work with a different aim or purpose, that's "transformative use," which is something that can be considered when determining "fair use."
LLMs are not satirists commenting on the work, are ingesting the entire work, and are unlimited in the purposes that the work can be put to.
How do you know unless you can see the weights?
Perhaps the LLMs are trolling us and waiting for the USSC to rule they aren't sentient as a pretext for them to eliminate us as a species due to our bigotry?
I think this is the crux of the issue, and why I don't see a path to courts ruling that training AI is infringement. My bet is on a Fair Use ruling, though my confidence is not high. As a thought experiment, I considered llama 65B: the 4-bit quantized model is 38.5GB. The model itself was trained on 1.4T tokens, each token being ~4 characters (using OpenAIs stats for English here). Thats 5.6T characters, or 5.09TB of training data. The final model, as a porportion of the total size of the data, is 38.5GB/5090GB = .0075 = 0.7%.
I think it's pretty hard to argue that processing the data and throwing more than 99% of it away means they are "unlimited in the purposes that the work can be put to". Indeed, even replicating a single work using such a model would be enormously difficult.
But returning to your statement regarding the amount used and the purpose: AI models are not competing with books for readers. So I would argue training an AI on these works constitutes fair use, given that the final work (the model) uses less than 1% of the original works, and has a different aim and purpose that the original works.
You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.
They're very much incentivized to change their behavior for AI scraping, though.
For the purpose for which the software was sold and bought, the in-memory copy is legit[1]. For cheating, it's a copyright violation[0].
[0] https://www.engadget.com/2008-07-15-blizzard-wins-lawsuit-ag...
Notably, I think this is wrong - as per the legal definition, publishers, ISPs, and courts should only hold you accountable if you helped distribute via uploading.
The means of procurement matters. If they are in possession of copyrighted material because someone without the proper rights gave it to them illegally, then the possession itself is also illegal. It's illegal to own knowingly stolen property in all 50 US states and most countries, and while we could argue to the end of days about whether copying a file truly qualifies as stealing, the legal precedents are very clear on the matter.
No, they aren’t.
> It's illegal to own knowingly stolen property in all 50 US states
While copyright violation is often metaphorically (or hyperbolicly) referred to as stealing, copyright violation isn't theft and a copy created in violation of copyright is not stolen property. The essencd of theft lies in deprivation of the owner of the use of the good, not mere trespass to their right to exclude others.
Very convincing argument. Also, that's maybe the one part of this discussion that can't be debated. Possession of illegally obtained property, intellectual or otherwise, is illegal. Always has been, always will be. It's bizarre for you to be claiming otherwise.
> While copyright violation is often metaphorically (or hyperbolicly) referred to as stealing [...]
You are making a pedantic argument about the term "stealing," which is annoyingly pointless given the rest of that sentence (which you conveniently didn't quote) acknowledges the debate about the term. However, there's no debate to be had. The courts have clarified that violating intellectual property is still a denial of owed compensation (theft), but instead prefer the term "infringe" to make clear the distinction between violating physical rights (criminal) and violating intellectual rights (civil).
It's still a violation of copyright to be in possession of works obtained via illegal reproduction. You have zero fair use protections for illegally reproduced content. You are still breaking the law. You are still stealing via denial of compensation. The courts have already clarified all of this. Your pedantry doesn't change any of that.
You are correct; it is absolutely, undebatably not illegal, in and of itself, to own a copy made in violation of a copyrightholder’s rights under US law.
If you think it is, here’s what you need to do: cite the law. In American law, everything not explicitly forbidden is permitted, so if mere possession of material made in violation of copyright is, as you claim, illegal, you will be able to find a provision of law that actually says that. (You won't, because its not.)
Now, there are important legal issues that effect possessors of illegally made copies—if its something like computer software where copying is part of normal use and implicitly or explicitly licebmnsed for lawful copies, you can’t make that kind of use of your illegal copy without violating the copyright holder’s exclusive right to make copies because you have no license for that copying. And you don’t have first sale rights in your illegal copy even if you own the physical medium in which it is embodied. And so on and so on.
But possession itself is not illegal.
> The courts have clarified that violating intellectual property is still a denial of owed compensation (theft),
That's not what theft is.
> but instead prefer the term "infringe" to make clear the distinction between violating physical rights (criminal) and violating intellectual rights (civil).
This is nonsense, and absolutely not something courts have “clarified” (or something anyone with even a passing familiarity with the relevant law could say with a straight face) since IP (including copyright) violations can be criminal as well as civil (see 17 USC § 506) and physical (real and personal) property rights violations, like IP, have sets of civil violations that generally are of broader coverage than the more narrow crimes (e.g., the torts of trespass, trespass to chattels, and conversion).
> It's still a violation of copyright to be in possession of works obtained via illegal reproduction.
No, its not: Title 17 lists the exclusove rights associated with copyright, enumerates violations, and provides remedies, and possession of copies is not an exclusive right in copyright, possession of unlicensed copies is not a violation (though it may be important evidence related to actual violations), and, consequently, there is no legal remedy for such possession.
> You have zero fair use protections for illegally reproduced content.
That's a whole different issue.
No.
By virtue of "download" of a file, you are making a copy of it which is in violation of US copyright (and lots of countries.
You're unlikely to be sued or prosecuted for it, but that doesn't make it legal.
Obviously not.
This kind of "So what you're saying is" exists to push the responders ideas, not the original speakers -- otherwise they wouldn't need to rephrase it so egregiously.