The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely
Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely
Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.
There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true
I am not personally affected because I don’t mind LLMs using my code and writing to learn. I have open source under MIT and similar licenses. I didn’t foresee LLMs learning from it, but it does feel like it’s in the spirit of what I intended.
Derivative work or transformative? It's not the same.
AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case
You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.
You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this
If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation
1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through
2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail
3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer
4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used
There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information
1. The plagiarism aspect, and that most of the training data was used without permission
2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)
Also, the restriction isn’t on a technology that could possibly reproduce something. It is on the act of using the technology to reproduce something.
> and yes that's 100% copyright infringement
Says what court of law?
I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.
The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.
>Says what court of law?
If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that
Only every movie piracy lawsuit ever. Nobody shares the original files after all, so every torrent is re-encoded in the way described.
Not only are you not winning this one but I'm gonna laugh at you the entire time.
Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)
Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for
There's also these parts:
> its use of pirated books to create such library does not constitute fair use.
Which indicates that there are tight bounds depending on the ethics of how the content was obtained
Similarly with the second one
>Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work.
We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards
>Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.”
Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings
I just pointed out none took place.
Anyway I'm not replying in this thread anymore.
A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.