We're literally talking in a thread about someone who committed suicide because the US government was hellbent on ruining his life with a felony conviction for piracy.
As we can see in the good article, the copyright system might almost have cost us a lot of AI capabilities. The damage it has done in cases where the lawyers got ahead of the builders is incalculable.
1. Unauthorised network access
2. Intent to distribute licensed material
The labs aren’t doing this. They are doing something similar to you and I downloading torrents. Look, I also think laws should apply somewhat equally to individuals and companies. But this is different.
According to the people trying to ruin his life. This is an accused thought crime, not something he actually did. Also, supposing it actually was something he did, your argument is that building a trillion dollar business on stolen licensed material is legally permissible but giving it away for free is worthy of your life being ruined. Wonderfully coherent world view, that is. Piracy is fine, but only if you hoard it to yourself and profit from it!
That doesn't seem different to me. Copyright, inherits.
If its fair use, like the companies currently claim, maybe it is different. But I suspect that the insane push to create a copyright carveout in countries around the world, is not unrelated to a judgement that it probably isn't fair use.
Song lyrics for example
It would be like if I sold you a pizza, but when you opened the box it contained all the source code for the latest GTA. The pizza wasn't copyrighted by anyone, but the what was inside the box was.
Strong IP advocates have argued for years that devices that can be used to infringe copyright are themselves infringement of copyright. So far that hasn’t held up to court analysis provided that device can be and is also used for non-infringing purposes. Given that so far the courts have found that training an AI model is sufficiently transformative to qualify as fair use, it doesn’t seem likely that distributing a model counts as distributing copyrighted material.
>training an AI model is sufficiently transformative to qualify as fair use
This is the key question and that courts have decided this way so far doesn't mean its the correct decision. If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.
> ...
> If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.
That may be so, IF you could actually do that. Yet in over 200 individual allegations in the Authors Guild vs. Open AI case complaint[1], not a single one of them alleges that you are able to do this. They allege that you could at one point get detailed verbatim quotations, but also note that the models have been explicitly blocked from doing this. Instead, the vast majority of the actual complaints in the case are about generating "summaries" that contain information not in other publicly available summaries, and generating "detailed outlines" of supposed future installments of the copyrighted works, using the characters and details of the story. In other words, all the actual alleged infringements are about either:
A) the copies made in order to train the model
B) infringing derivative works generated by prompting the AI
C) the verbatim copies made from sources to which OpenAI did not have rights to
You would think if any of the authors at all had been able to print out verbatim copies of their works without having to take knowing and direct action to circumvent the blocks in place to prevent that from happening, those would have been some of the top complaints in the case. The same held true for Bartz vs. Anthropic, where the judge even noted in his ruling that while generating verbatim copies might indeed be infringement, the plaintiffs never alleged that had happened or was possible.
And since Bartz vs. Anthropic has (reasonably IMO) found the training to be sufficiently transformative as to be fair use, the complaints for point A are unlikely to succeed here. The complaints for C almost certainly will succeed, for the same reasons they succeeded against Anthropic.
That leaves B. The questions would be:
1) Are such "detailed" summaries infringing just because they can include things other summaries have yet to include? Personally, I doubt they're going to get much traction here unless the courts split and they win on point A. The fact that other summarizers have left certain details out does not inherently make a new summary with other details an infringing work. If we imagine a world where Empire Strikes Back is a new movie, if none of the public reviews of the movie reveal the twist, but you can ask an AI model to summarize the movie and the AI model reveals the twist, that might be disappointing, but I don't think there's any argument to be made that it is copyright infringement.
2) Are speculative outlines of future unpublished work based on the information in a published work infringing just because they have been created?Is a model that CAN be used by a user to intentionally create an infringing derivative work itself an infringing product? Again without splitting the courts and winning on "training is infringing therefore all outputs are also infringing" I just don't see how they can win here. A speculative outline of future works is something people have been doing forever (see also any fan site on the internet). While attempting to publish that outline commercially or produce a new work from that outline might itself be infringement, that infringement is the result of explicit and knowing actions of the user akin to putting a book on a xerox machine and producing a cut and paste fan edit from the work. Again the xerox machine is not itself the infringement, and the individual page copies probably are also not infringement until they are used in a specifically infringing way.
3) Does the ability of the model to theoretically produce verbatim copies of the training material if OpenAI were to re-program the model to remove the blocks they have put in place to do that mean the models are themselves infringing. This is perhaps the most "up in the air" question of the 3, but the law generally doesn't award damages on the potential for copyright infringement, only on actual acts of infringement. Handbrake and various DVD copying tools do not ship with the keys necessary to defeat the DVD protection schemes, yet they know how to use those keys and accept such keys provided by the users. As far as I know, no court cases have been brought or succeeded against any distributors of DVD ripping software despite the fact that evading the "anti-infringement" blocks in the software is both trivial and exposed to the end user. Given that evading the "anti-infringement" blocks of OpenAI's models is neither trivial nor exposed to the end user, I'm fairly comfortable saying that again without splitting the courts and winning on point A, the authors guild isn't likely to win here either.
[1]: https://authorsguild.org/app/uploads/2023/12/Authors-Guild-O...
By comparison, the model isn't "watching" the show, as so many people are quick to point out that the "learning" analogy for what AIs are doing is flawed. There was never an intent by the creators that the show would be used to generate mathematical probabilities and weights in a statistical model and no one is deriving entertainment from making the statistical model. I suppose perhaps someone derives entertainment from AI training, but I suspect the number is small enough that "no one" is a reasonable approximation. So using the show to do so at least has an argument towards fair use. Or if the copy used for training was legally purchased, at least in the US it has the actual legal designation as fair use so far.
Don't get me wrong, I'm not saying that we should be returning to the days of the RIAA suing teenagers for their college education funds. But it seems pretty obvious that "pirating copyrighted material to explicitly use that material in the way that the creators of the material envisioned selling to you" is similar to, but arguably worse than "using copyrighted material (pirated or not) in a way not envisioned by the creator of that material to create a wholly different product". In both cases, the livelihood of the creator is possibly being affected, but one of them is a direct 1 for 1 loss of income while the other (again, if not specifically pirated) is an indirect impact.