It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.
Lesson in there about experts and politics.
> We document a tripling in the number of new books coming to market between late 2022 and late 2025 that mirrors the use of AI that we detect in new books. The effects of this influx on consumer welfare depend on the quality of the additional books. The average quality of new books has fallen with the LLM-induced influx, and books with detected AI are substantially worse than human-authored books, so that much of the new work is of little value to consumers. Still, the LLM influx has delivered some books in the middle range of the usage/quality distribution, and the LLM-era entry process delivered seven percent more consumer surplus from books than the pre-LLM process in 2025.
...
Moreover, the arrival of LLMs does not appear to have displaced activity by incumbent authors. Despite the controversy surrounding LLMs, their effect on book consumers – like other cost-reducing technological changes in the cultural industries – is positive. However, because the new books are mostly of low quality, the effects are modest
So not only are existing authors unharmed (because most of the new competition is slop) there is even a small improvement for consumers.
Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"
William Roper: "Yes, I’d cut down every law in England to do that!"
Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"
A lawless man can be struck down with force without persecution by law, because they are lawless.
Doesn't this define modern day police force theory?
If a lawless man robbed from a village, any of the villagers could do whatever they wanted to them in retribution. Maybe limited a bit by religion.
When an American becomes an outlaw, it doesn't mean the same thing as it did back then. There is no legal way for the government to deny someone their rights.
Ill gotten copyrighted material is illegal. What can be done with it after is a completely separate issue.
I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.
In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.
I suspect human writing could survive, but I don’t see any room for publishers.
No because we are people and the laws differ for people, corporations, and machines.
Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?
Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.
The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.
Which system of justice works like this? The law, uniformly applied, is a steamroller. That's why we have courts, to allow people to explain their actions (justify them).
Because the goal of laws is to improve human flourishing, not to be consistent. Laws are not strictly based on some sort of virtue ethics, they are often practical ways to accomplish the task of improving human lives. If having "applies to X but not Y and maybe Z depending on some criteria" accomplishes that then... that's the whole point.
> but they cannot be just, inherently
That's sort of an absurdly strong assertion. Why would exceptions not be "just"? "Killing someone is wrong, except in the case where it is strictly necessary to save lives in self defense" etc are generally consider just exceptions. This seems trivial. Very few people hold to an actual system of ethics that does not take context into account...
You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.
Yes, for some texts that's possible. But for the vast majority, it is not.
Or is it simply that the correct prompt hasn't been written for all possible cases?
I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.
It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.
Why is a vector database somehow completely different from maintaining a library of the text itself?
LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.
I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans
Here's a highly compressed representation of The Lord of The Rings (all three volumes):
1
Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no".
If I had compressed it by removing the letters x y and z, I'd agree with you that my "compressed" version is infringing.
So what we've got here is a spectrum with two ridiculous extremes, and a question: When has the artifact been compressed so heavily that it no longer infringes the copyright of the original?
I suggest "irretrievability" is a pretty good threshold for that question. Otherwise you're into "we know it infringes our copyright. Don't ask us to prove it, we just know it, ok?"
Given the sheer volume of text that an LLM gets trained on, and how small the output is, it seems obvious that 99% of it can no longer be recovered - the process is "lossy" to the point of irretrievability, and only a statistical smear is left behind. That's why I think only the copyright claims that can show infringement in court (Harpy Potter, et al.) have merit. And a court will still have to decide "how much is too much" but at least there's case law for that.
(Incidentally, I compressed the Mona Lisa to a single pixel. It was #3D3526).
I've tried my best to show where I think you're wrong. I think all that's left is arguing over the exact definitions of "recoverable" and "irretrievable". As I said, the courts will have to decide that.
Regardless, the argument that LLM output is or is not subject to copyright based on information theory is entirely defeated by what I've said.
Point an LLM at the conversation and ask it to ELI5 the competing arguments.
> Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.
To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.
My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.
(Edit: Think it was this one https://arxiv.org/abs/2601.02671)
That's one hell of a compression ratio, if it can do what you claim.
The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.
The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...
https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...
> had produced near-verbatim replicas
Since most users of LLMs are not using them to run around copyright and this copyright stuff is aggressively trained out of models or blocked whenever possible, the models themselves are still transformative works not intended to facilitate infringement.
That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.
https://arxiv.org/abs/2601.02671
The point is that for most texts, it is not possible. It's not able to recall what I wrote on Geocities in 1995, even though there's a good chance it was trained on it.
What if they retouch a photo you've taken as a political message? Does it matter if the do it with Sharpie, Photoshop, or by feeding it into an LLM?
When you publish, you give up some control over your work. Other people are allowed to do things with it and IMO it does not matter if it's in their head, on paper, or in a computer.
Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.
Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…
How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?
A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?
With open source, I should note, its remixing is in fact governed by copyright.
I think we’ve veered a little off the original thesis where we started which is an argument that copyright is limiting the rate or breadth or level of cultural artifacts. I have to say that I find it hard to imagine a meaningfully higher volume or level than we already see today. I mean are you worried that we’re stifling creativity? I think it is abundant and the evidence is all around us.
There might be something to that logic but art and literature doesn't obey rules like math and copyright exists to protect creators
So I’d be curious to hear about a counter example.
The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or whatever it is, is clearly ridiculous. 20 years before entering public domain seems pretty reasonable.
There's unfortunately also some pressure against people using legitimate public domain works. E.g. youtubers getting copyright strikes for playing public domain music because it's too similar to a specific copyrighted recording.
The problem is that there is also a lot of stuff that never got done because of copyright. And the extend to which works that were funded by exploiting copyright would not have been funded in any other way is also questionable.
Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.
Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.
That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.
We already have these - CD taxes, government grants funded by general taxes, GEMA in Germany, even TV licenses.
They all universally suck and are extremely unfair in who gets paid by them.
Yes, all the rich class propaganda being pushed by open source developers working on software in their free time.
Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.
If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.
There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.
There needs to be a way to support the creation of art.
A is not necessary for B if a single instance of B exists without A.
In colloquial terms it is frequently used to suggest a recommendation with an imperitive need.
Different things for the same word. Only the former can be used to determine if something is necessary, the latter is a subjective assertion and can have no proof either way.
It awards a few creators outsize rewards, but suppress creation of many more.
It does not petform the job that it is supposed to do.
Your entire thesis is false on its face. One does not need a big publisher to get published or make there work available. You've also offered no other alternative wherein the other works not sought by large publishers will somehow be afforded equivalent treatment so your proposition is just ridiculous if not outright ignorant.
Yeah, it's called "copyright."
Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
This is just bullshit and no one said it's the only incentive.
> There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
such as??
You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?
We don't need to imagine, this is how human society has worked for most of the run we have had.
A: rich people pay less % in taxes than wage workers, we should close the loopholes
B: but taxes are immoral to begin with
A: ok, but can we do something now about the unequal enforcement? Unrealized gains, tax havens, trusts, fake charities, etc?
B: well a society based on property rights… ackhully you should read this book by Mises/Rothbard/Rand
https://youtu.be/lh2__MN-FTU?si=LXIaljh__s8fD75l&t=1568
About 3 minutes of video worth watching.
For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.
This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.
The internet, in true internet fashion, still has the general logic level of a 15 year old.