If Meta (or anyone) had approached publishers with a “we want to buy one copy of every book you publish”, that doesn’t seem technical or business difficult.
Certainly Amazon would find that extremely easy.
This offhandedly seems to dismiss the cost of achieving legal clarity for using a book - a cost that will far eclipse the cost of the book itself.
In that light, it seems like an underweighted statement.
I'd like to see them try to argue Cartoon Network, LP v. CSC Holdings, Inc. applies to their corpus.
But the first one is a human using things. Its big guy vs little guy.
The prescident is there, google already "reads" every page in the internet and injests it into its systems and has for decades and has survived lawsuits to do so.
Peak was using MAI operating system directly by live booting it without their permission.
Antivirus and security companies don't need licenses to scan copyrighted materials to look for threats or vulnerabilities.
AI similarly is not executing, deploying, reselling or redistributing the copyrighted material. It's using the data to build a model. Security software distills the down data more, but it's still the same principle.
The million dollar question is if this counts.
IMO copyright law does not control what you can do with a book once you’ve bought a license, except for reproduction. It’s arguable that LLMs engage in illegal distribution, but that’s a totally different question from whether simple training is illegal even if the model is never made available to anyone.
Hypothetical: If the only way we could build AGI would be to somehow read everyone's brain at least once, would it be worth just ignoring everyone's wish regarding privacy one time to suck up this data and have AGI moving forward?
Would you trust a businessman on that?
The fact that this is an active question is depressing.
The suspicion that, if it were possible, some tech bro would absolutely do it (and smugly justify it to themselves using Rokkos Basalisk or something) makes me actually angry.
I get that you're just asking a hypothetical. If I asked "Hypothetical: what if we just killed all the technologists" you'd rightly see me as a horrible person.
Damn. This site and its people. What an experience.
I say it is no different than the people who are claiming they don't care. They absolutely do care, but at this point, saying "no" makes you the odd one with obviously something to hide, so they do this from a place of duress.
Unfortunately, I feel we are not too far from people finally snapping and going off the deep end because it's so pervasive and in-your-face that there is seemingly no escape left.
tbh human rights are all an illusion especially if you are at the bottom of society like me. no way I will survive so if a part of me survives as training data I guess better than nothing?
imo the only way this could happen is a global collaboration without telling anyone. the AGI would know everything about all humans but its existence has to be kept a secret at least for the first n generations so it will lead to life being gameified without anyone knowing it will be eugenics but on a global scale
so many will be culled but the AGI would know how to make it look normal to prevent resistance from forming a war here a war there, law passed here etc so copyright being ignored kind of makes sense
If it matched human intellectual productivity capacity, that ensures that human intelligence will no longer get you more money than it takes to run some GPUs, so it would presumably become optional.
But it’s not at all a similar dilemma to “should we allow the IP empire-building of the 1900’s to claim ownership over the concept of learning from copyrighted material”.
As I wrote it out, I didn't know what I thought either.
But now some sleep later, I feel like the answer is pretty clearly "No, not worth it", at least from myself.
Our exclusive control over access to our mind, is our essential form of self-determination, and what it means to be an individual in society. Cross that boundary (forcefully no less) and it's probably one of the worst ways you could violate a human.
Besides, I'm personally not hugely into the whole "aggregate benefits could outweigh individual harms" mindset utilitarians tend to employ, and feels like it misses thinking about the humans involved.
Anyways, sorry if the question upset some people, it wasn't meant to paint any specific picture but a thought experiment more or less, as we inch closer to scarier stuff being possible.
Sure, if we all get a stake of ownership in it.
If some private company is going to be the main beneficiary, no, and hell no.
But we do, in the sense that benefits flow to the prompter, not the AI developers. The person comes with a problem, AI generates responses, they stand to benefit because it was their problem, the AI provider makes cents per million tokens.
AI benefits follow the owners of problems. That person might have started a projct or taken a turn in their life as a result, the benefit is unquantifiable.
LLMs are like Linux, they empower everyone, and benefits are tied to usage not development.
The price will be ratcheted up, such that the majority of the economic surplus will go to the owner of AGI - with pricing tiered to the query you're asking it. The more economic utility the user will derive from making the query, the more the AGI's owner will charge them.
So in this scenario I could see it become necessary from a military perspective.
> trained the model on synthetic data
You get knowledge collapse [1] this way.But even though there are counter-examples (e.g. learning Go from self-play based only on the rules), IMO — based on how often people were already doing this with buggy human-written software[0][1] — most people don't think about this in the right way and will therefore treat these things as magical oracles when they shouldn't.
No silver bullets.
[0] https://en.wikipedia.org/wiki/British_Post_Office_scandal
I don't understand.
Facebook and Google spend billions on training LLMs. Buying 1M ebooks at $50 each would only cost $50M.
They also have >100k engineers. If they shard the ebook buying across their workforce, everyone has to buy 10 ebooks, which will be done in 10 minutes.
Most universities have had their own corpora to work with, for example: the Brown Corpus, the British National Corpus, and the Penn Treebank.
Similar corpora exist for images and video, usually created in association with national broadcasting services. News video is particularly interesting because they usually contain closed captions, which allows for multi-modal training.
youtube, etc classifiers definitely do read others material though.
LLMs are much smaller than their training sets, there is no space to memorize the training data. They might memorize small snippets but never full books. They are the worst infringement tools ever made - why replicate Harry Potter by LLM, it's show, expensive and lossy, when you could download the book so much easier.
A second argument is that using the LLM blends a new intent into the process, that of the prompter. This can render the outputs transformative. And most LLM interactions are one-time use, like a scratch pad not like a finished work.
How many bits of entropy in Harry Potter?
How many bits of entropy in a lossy-compressed abridgement that is nevertheless enough, when reconstituted, to constitute a copyright infringement of Harry Potter?
The latter is absolutely small enough to fit in an LLM, although how close it would get to the original work is debatable. The question is whether copyright is violated:
1) inherently by the model operator, during the training.
2) by the model/model owner, as part of the generation.
3) by the user, in making the model so so and then reproducing the result.
1) straight up copying. download a bunch of copyrighted stuff -> making a copy. no way out of this one.
2) a derivative work can be/is being generated here. very grey area — what counts as a “derivative” work? read about robin thicke blurred lines court case for a rollercoaster of a time about derivative musical works.
3) making the model so so? do you mean getting an output and user copying the result? that’s copying the derivative work, which, depends on whatever copyright agreement happens once a derivative work claim is sorted out.
that’s based on my 5 years of music copyright experience, although it was about ten years ago now so might be some stuff i’ve got wrong there.
If abstract ideas were protectable what would stop a LLM from learning not from the original source but from social commentary and follow up works? We can't ask people not to reproduce ideas they read about. But on the other hand, protecting abstractions would kneecap creativity both in humans and AI.
2) Agreed! Where this becomes interesting with LLMs is that, as with people, they can have the capacity to produce a derivative work even without having seen the original.
For example, an LLM that had "read" enough reviews of Harry Potter might be able to produce a reasonable stab at the book (at least enough so for the law to consider it a derivative) without ever having consumed the work itself or direct derivatives.
3) It's more of a tool-use and intent argument. One might make the argument that an LLM is a machine, not a set of content/data, and that the liability for what it does sits firmly with the user/operator, not those who made it. If I use a typewriter to copy Harry Potter - or a weapon to hurt or kill someone - in neither case does the machine or its maker have any liability there.
I think this is a dangerous road with little upside for anyone outside of IP aggregators.
The real question is - does copyright grant the authors' the right to control if their work is used for LLM training?
Its not obvious what the answer is.
If authors don't have that right to begin with then there is no way amazon could buy it off them.
Let me put that into perspective:
- Googling "how many books exist" gives me ~150 million, no idea how accurate but let's use that. - Meta had a net profit of ~40 billion USD in 2023. - That could be an potential investment of ~250 USD per book acquisition.
That sounds like a ludicrously high budget to me. So yeah, Meta could very well pay. It would still not be ethical to slurp up all that content into their slop machine but there is zero justification to pirate it all, with these kinds of money involved.
Anyway the problem is not money it's technical feasibility and timelines.
You clearly don't think LLMs have any value though, so W/E
Absolutely no way. Yup.
> Buying millions of ebooks online would take a lot of effort, downloading data from publishers isn't a thing that can be done efficiently
Oh no, it takes effort and can't be done efficiently, poor Google!
How can this possibly be an excuse? This is such a detached SV Zuckerberg "move fast and break things"-like take.
There's just no way for a lot of people to efficiently get out of poverty without kidnapping and ransoming someone, it would take a lot of effort.
Edit: Changed it just for you
Uh, pardon? For a mere $10MM, you can get almost all of the Taylor & Francis' catalogue. They'll pressure their authors to finish their books early for free [0].
I think you can obtain all the training material for a mere rounding error in your books, if you're Meta, or Microsoft, or similar.
Well, the authors will not be notified, compensated, or their idea on the matter won't be asked anyway, but this is "all for capit^H^H^H^H^H research".
[0]: https://mathstodon.xyz/@johncarlosbaez/113221679747517432
This is not copyright as we know it. Copyright protects against copying, not accessing data. You can still compile statistics off data you don't own. The models are like a compressed version of the originals, so compressed you can't retrieve more than a few snippets of original text. Newer model train on filtered synthetic text, which is one step removed from the protected expression in the copyrighted works. Should abstractions be protected by copyright?
was thinking if i train my model on my private docs for instance finance how does one prevent the model from sharing that data verbatim
1. Sam Altman was removed from OpenAI due to his ties to a Chinese cyber army group.
2.OpenAI had been using data from D2 to train its AI models.
3. The Chinese government raised concerns about this arrangement with the Biden administration.
4. The NSA launched an investigation, which confirmed OpenAI's use of D2 data.
5. Satya Nadella ordered Altman's removal after being informed of the findings.
6. Altman refused to disclose this information to the OpenAI board.
Source: https://www.teamblind.com/post/I-know-why-Sam-Altman-was-fir...
I guess Sam then hired top NSA guy to buy favor with the natsec community.
I wonder who protects Sam up top and why aren't they protecting Zuck? Is Sam just better at bribes and manipulation?
"It would take a lot of effort to do it legally" is a pathetic excuse for a company of Meta's size.
Defending themselves with technicalities and expensive lawyers may be financially viable.
Zero ethics but what would we expect from them?
While it's plausible someone downloaded a bunch of torrents and tossed them in the training directory...again, under who's authority? Like if this happened it would be one overzealous data scientist potentially. Hardly "them".
People lean on collective pronouns to avoid actually thinking about the mechanics of human enterprise and you get extremely absurd conclusions.
(it is not outside the bounds of thinkable that an org could in fact have a very bad culture like this, but I know people who work for Meta, who also have law degrees - they're well aware of the potential problems).
These newly unredacted documents reveal exchanges between Meta employees unearthed in the discovery process, like a Meta engineer telling a colleague that they hesitated to access LibGen data because “torrenting from a [Meta-owned] corporate laptop doesn’t feel right ”. They also allege that internal discussions about using LibGen data were escalated to Meta CEO Mark Zuckerberg (referred to as "MZ" in the memo handed over during discovery) and that Meta's AI team was "approved to use" the pirated material.
https://www.wired.com/story/new-documents-unredacted-meta-co...I totally agree. But since when has that stopped companies like Meta. These big companies are built on breaking/skirting the rules.
They could also simply buy controlling stakes in publishers. For scale comparison, Meta is spending upwards $30B per year on AI, and the recent sale of Simon & Schuster that didn't go through was for a mere $2.2B.
Surely the author only licenses the copyright to the publisher for hardback, paperback and ebook, with an agreed-upon royalty rate?
And if someone wants the rights for some other purpose, like translation or making a film or producing merchandise, they have to go to the author and negotiate additional rights?
Meta giving a few billion to authors would probably mend a lot of hearts, though.