But I hear you. One of my biggest tells that someone can't be reasoned with is when they resort to whataboutism without any consideration for how 2 situations can actually be different even if there is some commonality. It's a powerful bad faith argument technique. When that style of argument comes up I nod my head and walk away. Some people are just doomed.
Don't they? They release the llama model weights, they do things like this:
https://www.opencompute.org/wiki/Open_Rack/SpecsAndDesigns
They also make significant contributions to Linux and are the originators of popular open source projects like zstd and React.
They make their money from selling ads, not selling licenses.
What do you think the outcome of tightening fair use is going to be? Do you think its going to be most effectual against these big evil AI companies we are meant to fear? Or is it going to end up putting more individual creators on the end of Disneys pitchforks?
Like if you support creating a gun to kill a monster, that's great. But you need to understand that weapons rarely only target the person you want them to. And its unlikely that any bill that specifically targets a certain size or profit margin is going to make it all the way into law without being generalised to the approval of large IP holders.
Its much much (much) better to look at this as an opportunity to erode IP laws for everyone, than to make them worse and hope that your particular enemies are the only ones that are affected.
>That doesn’t mean I’m behind industrialized narcotic production on such a huge scale that it that it starts to distort the economy, and companies looking for new ways to add methamphetamine to every goddamn product.
Thats such a non sequitur. This isnt a weed legalisation argument, its "Do we make IP worse for everyone, because you dont like some people benefiting from fair use".
Meta used allegedly stolen copyrighted materials to train a model they shared for free with the whole world. Is this a recreational use?
It is the same playbook everytime. We dont have to be naive and pretend meta is doing something for other peoples benefit.
Are you unable to access this page?
https://www.llama.com/llama-downloads/
Or this one?
https://lmstudio.ai/models/meta/llama-3.3-70b
>It is use to build a monopoly
How?
>We dont have to be naive and pretend meta is doing something for other peoples benefit.
Meta benefits from the current war of open model competition, but we also benefit from it. In particular, participating in all this makes it hard for them to pull the ladder up when the market changes. They will have to justify why whatever new hotness is better than these existing models already on our hard drives.
'"They then copied those stolen fruits"
How are these fruits "stolen" if they still have what was allegedley stolen?
Dowling v. United States, 473 U.S. 207 (1985): The Supreme Court ruled that the unauthorized sale of phonorecords of copyrighted musical compositions does not constitute "stolen, converted or taken by fraud" goods under the National Stolen Property Act
And even if, arguendo, sure its stolen. The purpose of copyright is to "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries"
And you would be hard pressed to prove that LLM's haven't advanced the arts and sciences, so at bare minimum transformative, ie fair use.'
Also social media profile pics. Great way to get faces for deep fake ads. Most people are just 1 phone call away from being voice cloned. Our likeness isn't all that important either if you think about it.
Maybe meta will clone your writing style and sign into your meta account and message your friends telling them about this awesome new product. Meta owns the account and you uploaded data to it.
Don't be naïve. A corporation would tear the flesh from your body if it meant a better quarterly earnings report.
If you write a book and I take it and embed its knowledge into my product that is so pervasive that no one needs to buy your book any more (and I don't even credit you so no one knows where that knowledge came from), to you really still have what was stolen? And I didn't even buy a copy of your book to copy it.
> Theft [...] is the act of taking another person's property or services without that person's permission or consent with the intent to deprive the rightful owner of it. --- https://en.wikipedia.org/wiki/Stealing
maybe you should look up the definition of property, which is a set of legally recognized rights over a thing, typically including:
* possession (what you're focusing on)
* use
* exclusion
* transfer
The last 3 seem like they have been breached, in legally that's theft.
Getting punched in the face also violates rights, yet isn't murder. Murder is specifically about dying.
With theft, the entire damage is the deprivation. It could be an heirloom or some other object that may have been entrusted to you, something that can never be replaced, memorabilia of loved ones. Something that you may have needed in your posession to survive (e.g. a car to go to your job).
With a given copyright violation, the damage is that maybe[1] you made less profit than you could have. The potential for profit is not property. Profit isn't guaranteed.
[1] The loss is not certain, because there's no guarantee that the ones consuming the copyrighted content could have even afforded it.
People seem to think what ai is today is theft. If enough people agree, it will be theft. Big companies dont like this and push the other way. An objectiveness doesnt exist here. It is too wiggly
This can even extend to stealing physical property.
Depending on local laws, stealing a car may not actually be theft if the defendent can prove they intended to return it before the owner got home from work, though it would certainly be considered theft in the colloquial sense of the term, and they would still be guilty of a lesser offense like civil and/or criminal conversion.
I doubt there's even one place where the law works like that.
In a lot of places, that's how it works. A key element of theft is the intent to permanently deprive someone of property.
This is why joyriding isn't classified as auto theft and is instead a lesser offense. It's because joyriding is an intent to temporarily deprive, while GTA is an intent to permanently deprive.
In some jxns (the UK is one), there is a tort called trespass to goods, and an example of this would be "stealing" someone's property to deliver to another location for them to use there. The tort of conversion is similar: interference with someone's property right to treat it as your own (silent as to length of time).
This isn't a court of law. We don't have to talk like lawyers. If you replaced "theft" with "copyright infringement" in the comment you had such a problem with, what meaningfully changes besides we all have about five additional brain cells?
The obvious difference that copyright is subject to fair use and various other limitations that personal property isn't.
The trouble with this analogy is that it proves too much.
Suppose you write a book, and so does someone else, but they have better marketing than you and then people in the market for that genre buy theirs instead of yours. Let's even stipulate that the existence of their book actually lowers your sales, because people who want that kind of book already bought theirs by the time they find out about yours and then some people don't have time to read or can't afford to buy both.
Notice that we haven't yet said a word about the contents of either book. They could be completely independent and they've never even heard of you or your book -- they "didn't even buy a copy of your book to copy it". All we know is that they're the same genre and the existence of theirs is costing you sales. By that logic all competition would thereby be "stealing", and that can't be right.
Which implies that you don't have a property right to the customers.
I'm not sure how this plays out legally, but it certainly seems unethical
Would it be fair for Greece to do retroactive term extensions all the way back to Plato and then sue anyone who copies the idea of having a university or uses the Platonic solids or distributes religious texts that incorporate the dualistic theory of the soul?
Letting companies train LLMs on the "classics" is very different to training on contemporary media where the creator still depends on it.
The attempt to distinguish them is through copying, but that's the part that isn't depriving anyone of anything.
People on this thread need to focus on what "derivative" and "fair use" mean and understand both are measured on a somewhat fuzzy spectrum, subject to interpretation.
In a perfectly fair world AIs/MLs could vacuum up all human knowledge, fair and square. (In an ideal world, they would do that adhering to polite opt-in/opt-out agreements with copyright holders. We can dream). Input isn't theft.
On output, two magic genies would stand at the gate, the Derivative Genie and Fair Use Genie and review anything spat out by the AI/ML. If it crossed agreed upon thresholds the Genies would bar the gates and issue a stern warning to prompt again (or maybe the AL/ML would auto-adjust the prompt and try again).
So, if your prompt asked for a 300-word poem about thrash metal mosh pit dancing and it spat out a poem where 85% of it match one of the handful of available mosh pit poems in its database, the Derivative Demon would block the output and raise an alarm.
On the other hand, if you asked for a line by line analysis of a famous mosh pit dancing poem (by name) or maybe asked for a satirical spoof of said poem, the Fair Use Demon would overrule the Derivative Demon and give the output a pass.
That's as fair as this could get, especially if you add one more thing: An Appeals Court (maybe corporate, maybe 3rd party, maybe state run) with a Settlement Pool. If a copyright holder could prove the Genies let pass something they shouldn't, the AL/ML would fix that. If real damage is done, the creator would get a settlement from the pool.
The point is that the Input Genie is out of the bottle. Creators just look foolish trying to squeeze it back in. Better, they should focus on making the output Genies and the Appeals process as effective and fair as possible for everyone.
Facts are not copyrightable. Only your particular way of expressing those facts is copyrightable.
I wouldn't even go that far. Its an entirely new product. Its like the guy who sold you the keyboard demanding royalties for the software you built.
That the person who wrote the book couldn't predict a new use case for the book in training LLMs, is irrelevant. The book isn't in the LLM. Its not being sold with the LLM. Its one of billions of tools used to create the LLM.
People try and sell this as the AI companies extracting value from the poor little IP holders like Disney. Its maddening. That content is your cultural heritage. It already belongs to you, just some idiot has been granted a lifetime of exclusive exploitation. An LLM is trained on data you already own. Disney et al wants to exploit the new technology to extract even more money out of stuff created often decades ago.
At absolute worst its reverse engineering, which was supposed to be fair use protected in the US but apparently that's been somewhat eroded.
An LLM is essentially a lossy compression of the training data. The book absolutely is in there, it’s just mangled to the point of unrecognizability.
When large quantities of source material are replicable by prompting its a bug not a feature.
>The LLM Wouldnt be here without the copyrighted works
Google wouldn't be here if it hadn't scraped every copyrighted website and used them to form a searchable graph of the internet but we only hear complaints about them when they reproduce entire news articles.
What makes you think you are entitled to tell people what they can and cant do with data they purchased (or otherwise acquired) from you. Extremely honest question. I just cant put myself in your shoes.
Like if I had written anything useful I would be overwhelmingly flattered that my content be considered so worthy for inclusion.
Your profile suggests that you are a philosopher. Did you get into philosophy hoping to exploit the publishing industry to the extent that you can squeeze every cent out of your thoughts, and deny their potential uses downstream?
Its actually crazy how bad things are, I am usually keen on capitalism and exclusivity, but the whole thing with LLMs, I see people pushing hard to tighten the grip of intellectual property. I see people making 50 cents a month on Kindle Unlimited suddenly shocked that someones LLM generated output might be ever so slightly influenced by weights ever so slightly influenced by their work, seemingly thinking they might get some big payday out of it.
Give me a tiny little wedge of understanding of your thought process. Your book is right now, doing a greater social good on your behalf than me running around and removing all the trash from my neighborhood, and the benefits of that social good are going to accrue long after you and I are gone. Your work is now going to live on, in a very tiny way, in these systems forever. I am honestly envious.
If anything, I would be trying to get bad writing removed from LLM training data. Things that I dont want to influence others. But as a potentially honest promoter of your work, you want it removed?
Whats the number? If not 1:1 exactly what you charge for the book, what do you think the proper compensation for slightly influencing training weights you should receive?
Hundreds of years of copyright law. I bought a copy of Windows, but I’m not allowed to modify that data with a cracker and sell a bootleg DVD of it.
I should edit to clarify that I’m not a big fan of Lars Ulrich or Disney, but I don’t think we’re going to get a win here for the recreational IP pirates. What’s more likely is that we’ll end up with some Frankenstein law that favors both Mikey Mouse and OpenAI, and you and I will neither get free movies nor the ability to earn a living off of our creative labor.
But in abstract you should absolutely be able to modify and sell windows.
And before you give me an analogy about how someone could listen to Pink Floyd and then produce works inspired by their influence yada yada: Someone is a human being with human rights, and if we're going to start pretending that training an LLM is in any way analogous to human consumption and creativity, and not an industrial process that encodes input data into a digital artifact, then let's start by saying LLMs have human rights and cannot be owned by a company that charges for access to them.
Yep and so far it looks like the issue with the meta case is they didnt pay for the book. Not that they used it in training data.
>in the most charitable interpretation, using them to create derivative works.
Yeah in the same way I use a hammer to create a derivative table.
>Someone is a human being with human rights, and if we're going to start pretending that training an LLM is in any way analogous to human consumption and creativity.
I dont care about that. Its simply a tool being built using existing tools. Like using a jigsaw to make a step ladder.
Let's not sane-wash what they did here, they didn't just 'forgot to pay for the books', they deliberately and illegally downloaded and used material that wasn't theirs to use.
If you or I did that, we would be jailed or sued into destitution. In a fair world we either should change copyright laws (allowing for anyone to freely pirate all media), or Zuckerberg needs to go to jail.
Yes. Forgot is your word.
But lets face it, there wouldn't be a case to answer for if they had paid retail for each book, torn them up and scanned them and trained on that data.
>Zuckerberg needs to go to jail.
I am comfortable with that but would prefer updating copyright.
It’s called a copyright notice. Same as a license. If you’re running a commercial business you can’t legally just take that piece of work and reuse it. Pick any book off your shelf and pretty well every one of them will have words to the effect of:
All rights reserved. No part of this publication may be reproduced, distributed, or transmitted in any form or by any means, including photocopying, recording, or other electronic or mechanical methods, without the prior written permission of the publisher, except in the case of brief quotations embodied in critical reviews and certain other noncommercial uses permitted by copyright law. For permission requests, write to the publisher, addressed "Attention: Permissions Coordinator," at the address below.
Same as every piece of commercial software has a license which has to be abided by. Same as use of Meta’s service has terms and conditions which HAVE to be agreed to.
So yeah they’re free to break that license but they’re also free to be sued by IP holders for breaking it at scale.
The items they call out around training the models (and attempting to claim that each subsequent model generation should count as an additional instance of infringement) seem far less grounded in the current court interpretations of AI training.
I am not a fan of US copyright law, but if I torrented millions of books, I would be facing a felony charge in criminal court and a (with statutory damages as high as $150,000 per title in cases of willful infringement) multi-billion dollar lawsuit in civil court.
In my opinion, this has nothing to do with whether or not AI training is transformative and this fair use, and everything to do with whether or not the laws apply to everyone equally. If Facebook isn't forced to pay billions and elect a sacrificial executive to serve prison time, then I will remain angry.
That is not what this case is about. It is more about the illegal violation and piracy of copyrighted content done by Meta for commercial use and Zuck knew they were doing it.
Why did Anthropic settle [0] with a multi-billion dollar payout to authors after commercializing their LLMs that was trained off of copyrighted content that was illegally obtained and kept without the authors permission?
There's a reason why they (Anthropic) did not want it to go to trial. (Anthropic knew they would lose and it would completely bankrupt them in the hundreds of billions.)
AI boosters will do anything to justify the mass piracy and illegal obtainment of copyrighted material for commercial use (not research) which that is not fair use in the US. There is no debate on this. [0]
[0] https://images.assettype.com/theleaflet/2025-09-27/mnuaifvw/...
The original work is not replicated identically, why would we replicate a work when it can be more easily seen in original or replaced with an alternative options online. We use AI to produce new outputs to new situations. We already have had drives and networking for plain copying.
Or anything to defend on Meta. If they go out of business, humanity profits.
Elsevier at least works within the (admittedly broken) system, Meta does not.
Not even going to all GPL stuff, that in a better world should have screwed all the slop companies