I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
100% of the rulings agree with me.
The piracy is not in question. It is unarguably copyright violation.
But that's not what anyone means in this context. Training is what everyone means.
> The law is the law, there can't be different law for corporations with billions in backing.
I didn't say otherwise. That's a straw man.
No.
https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
This order grants summary judgment for Anthropic that the training use was a fair use.
And, it grants that the print-to-digital format change was a fair use for a different reason. But it
denies summary judgment for Anthropic that the pirated library copies must be treated as
training copies.
The document you linked was written a month before the Bartz decision was reached. It's also worth noting even the document you linked says this in its conclusion: Various uses of copyrighted works in AI training are likely to be transformative. The
extent to which they are fair, however, will depend on what works were used, from what
source, for what purpose, and with what controls on the outputs—all of which can affect the
market.
[1]: https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...That is not what the text you quoted is saying. It says that the output may be transformative, which is one criteria, but the other criteria depends on how it is used and what the source is.
Because torrenting includes redistribution, since you also upload to other peers
It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling
Exactly analogous to a human reading the material.
Yes - but some people can.
Are they criminals for reading books?
You're not making sense.
The problem is the reciting. Not the reading. And hence, not the training.
Whether this is right or not is a seperate question.
That is not the situation we are discussing. No one is arguing that what you describe is infringement.
What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.
Current copyright laws are simply not prepared for this unprecedented use.
Reading a book embeds its contents into your brain. And yet, that is considered fair use.
I agree, the analogies are meaningless in a legal context. In the legal context, the courts disagree with you.
It will give you the complete song lyrics which are under copyright.
The infringement occurs at the time of copy, not at the time of training.
What you can't do is distribute the lyrics without the author's permission.
When you ask DeepSeek to retrieve the lyrics (and it does so), the real question is this: Is DeepSeek merely acting as an intermediary/ISP according to the DMCA (in which case they'd be protected under the safe harbor clauses) or are they illegally redistributing the lyrics without the author's permission?
Whether or not you own a copy of the lyrics is irrelevant from a legal perspective in this scenario.
My guess: If they just retrieved the lyrics from some website and delivered them to you (because you asked), they're an ISP. However, if they pulled them out of their own database, they're violating copyright.
LLMs are far too lossy to be able to store such lyrics in their entirety. In fact, they're not even "lossy" since they're not even trying to record such information. They're just weights for how likely it is that any given word will come after another.
And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?
Excusing illegal behavior creates a moral hazard and over time leads to the kind of amoral and corrupt government and business leaders we have today. There's a pervasive sense in society that money can buy you out of any consequences and morality is basically irrelevant, every little thing like this is how it happens. It's the same dynamic as SF failing to police minor shoplifting for years, which led to stores closing and metal bars everywhere that we have today. Even then we actually do punish a few shoplifters for stealing $100 worth of goods, but we don't punish these companies for stealing billions worth?
Of course it doesn't. No one is arguing that.
We are arguing the bigger issue of whether training is an infringement.
> The only reason they haven't been punished
The reason they haven't been punished is because copyright infringement is not a criminal offense, it's civil. And the labs are settling those cases.
Not by the courts it's not.
I think it might be against the terms of service.
That has yet to be decided in court. At best it's a violation of the terms of service, but the only restitution for such a violation is termination of service and by then it's too late.
Remember: Distillation is literally just giving the AI a prompt and seeing what it spits out. If that were illegal, we'd all be guilty whenever we used AI.
Examining your competitors doing business is as old as business.
Judging by how AI threads look like for the past year, they were absolutely right to be worried.
> largest copyright theft operation in human history
In fact, you're doing exactly that right here.
You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.
Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
[1] https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s...
You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.
Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve of libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.