Zuckerberg approved training Llama on LibGen [pdf]
storage.courtlistener.com
storage.courtlistener.com
Or maybe we will come into the conclusion that all this works only if there's no such thing as IP, reset the playing field for everyone and if anyone wants to make money will have to actually work for it every single time. IIRC that's what's happening in China and its how they surpassed US in innovation.
Technically, that's a deregulation - just not the kind of deregulation the big tech is pushing for. Maybe the next time there's a graph showing how regulations made EU lag behind, add the graph of China too to spice things up.
With so many technical people out of work and promises of make the employed ones obsolete too, it can be a good idea to let people build thing instead of unfairly concentrating even more power onto kleptocratic entities.
Even in the 18th century, the French aristocracy mostly cruised through the Revolution from afar, surviving with fortunes largely intact to this day [1]. If the fork is UBI or guillotine, the selfish move by the private-jetting billionaire class—personally and financially more mobile and global than the French aristocracy ever was—is the latter.
> if there's no such thing as IP, reset the playing field for everyone
Your thesis is letting Altman, Zuckerberg and Musk have free rein would decrease inequality?
> IIRC that's what's happening in China
Not really [2].
[1] https://www.bbc.com/news/magazine-37655777
[2] https://www.chinaiplawupdate.com/2023/08/china-prosecutes-11...
Americans are largely not for a revolution because most of us aren’t idiots. There is idle chatter of a civil war, but that’s again (a) bluster (not that this can’t take on a life of its own) and (b) about consolidating control versus wholesale rebuilding the American class structure.
Anyone advocating for the first (as a popular revolt) thinking it wouldn’t result in the second isn’t thinking realistically.
There are plenty of examples in modern Europe where revolutions and regime changes didn't involve a civil war.
Where internal power structures were preserved (or where the society was restructured under occupation), yes.
No. See for example Spain's or Portugal's transition from autocracy to democracy. The latter involved a military coup and exile of it's former dictator.
The American Revolution was one of American elites overthrowing their overseers. It worked and was not super disruptive because power (and class) structures were preserved. From the states through to the system of law and the people in power. (We also didn’t do any mass or political executions.)
We have zero historical or contemporary precedent for this, and strong incentives for everyone else in the world to not play along. (As they did in sheltering the French aristocracy.)
In a hypothetical American revolution, foreign powers would be looking for their slice of the pie. To think through this dispassionately, imagine civil war breaking out in Russia or China. A second American revolution à la the first would put today’s billionaires and political elite in a room to draft a new constitution to their liking.
> Criminal trademark infringement made up the majority of IP crimes with 10,384 people prosecuted accounting for 88.9% of the total.
Trademark infringement is of a completely different character from copyright.
Trademark infringement is pure fraud and lying.
Take out trademark infringement, and you have only 1 prosecution per year per 700,000 people.
What is it in America? Did we even have a single criminal non-trademark IP prosecution in 2024?
There is a reason everyone with over 130 IQ wants to work for them rather than starting their own companies.
We can’t protect IPs only when that benefits big corps. We should protect them always or accept that the world is better if we go in another direction, changing the rules for everybody.
- of course exact reproduction of protected content is a no-no
- but learning is ok, as long as it is transformative. User prompts and responses are pushing the model outside its training distribution anyway - users add their own intent, making usage transformative
- when LLMs synthesize from multiple sources, the result is transformative
- if you try to protect expression it is meaningless now, but if you protect abstract ideas it kneecaps creativity
- the problems of copyright started with the apparition of internet, not with AI
- revenues from royalty are almost zero today, as each new content competes against an unbounded list of other works that have been accumulating for decades online
- because royalties are shit, creatives now focus on ads, and this leads to enshittification, attention grabbing junk everywhere, attention is scarce content is post-scarcity
- we actually like interactive participation more than passive consumption; we now edit Wikipedia, contribute to open source, have papers published for free on arXiv, use social networks where our comments are shared with the world, play games instead of reading books - it is another age, the interactive age
- AI is actually more than an infringement tool, it is useful for many legit purposes
- and AI is the worst possible infringement tool, it can hallucinate details, get thins wrong; By comparison copying is free and easy and precise to the letter
So the idea that training is infringement is pretty abusive, it tries to make copyright be about abstractions which is wrong. We can't return to 1990s, so we have to live with its demise. It's been dying for 3 decades already.
Even if you forbid AI from training on copyrighted works, people are going to comment about them online, and the model will pick up the ideas. There is no way to protect ideas from spreading and reaching AIs.
Is there a reason a human can't torrent movies and say "But I'm just learning from them"?
There's a reason why every vassal with a sizeable estate wants to be in the King's court rather than starting their own country.
Realistically billionaires are using racist and homophobic populism as a way to direct working class energy away from wealth inequality. Making people think "woke" is the reason why the earth is on fire and they can't have health insurance.
We saw a bit of that with Covid cheques.
I'm not against the idea of UBI, I just see the landlords eating it up like they do with peoples wages.
Though the state would have to make sure the person receiving the benefit actually exists, is still alive, etc.
The very nature of Cantillon is unequally obtained new money, whereas UBI is universal. Any effect it has would be related to the poorest/neediest spenders now purchasing the sort of goods they do (and, realistically, no increase in spending by the richest). You might see increased consumption in neighborhoods/regions with high concentrations of poor, too.
The better fit for "UBI creates an economic problem" seems to be pricing stickiness. The above commenter focused on controlling general inflation through monetary/fiscal policy (keeping money supply stable, using tax mechanisms), but didn't actually address the concern about producers simply raising prices to absorb the UBI.
It's gonna be interesting that's for sure.
Even assuming this scaremongering scenario, the world would be in a far better place if society assured everyone would be guaranteed a certain income.
Also, the scenario that supports the hypothesis of higher inflation is that more people in society are suddenly able to afford goods and services that were out of their reach without UBI. Can anyone actually put to words why that is undesirable?
- If everyone suddenly has more money (say $2 more per day)
- And milk is a basic necessity
- The milk seller knows everyone needs milk and now has $2 more to spend
- They can gradually raise the price of milk by close to $2
- Consumers must still buy milk at the higher price
- The intended benefit of the extra $2 is effectively captured by the milk seller
The increases in general purchasing power can be absorbed by suppliers of essential goods. If you have just excess discretionary income in the general case, then non-essential goods can bump in price, too.
For the sake of argument, imagine UBI provides everyone with a million dollars a year. That doesn’t make everyone a millionaire. It just makes everyone’s money less valuable.
It’s no different on a smaller scale.
That slacker is already getting high and playing on Xbox. With UBI they will have less worries about staying alive and the opportunity to try things to get more money. UBI is a great insentive for people to try new things without there being a financial risk of you losing your income. Just check the trials and their results - people are more productive and happy in general.
The trained models are trillionths the size of their training sets. There is no archive of copied data in them.
Acquiring and using works without such license is just piracy. Whatever your stand on piracy is, most individuals and businesses are not free to incorporate it into their projects. Normal people have faced significant penalties for piracy, and concientious business operators avoid it.
Sure would be disappointing to all those people if there were suddenly a ruling that said "well, but it's okay that these guys did it because they're filthy rich and went real hard with it"
Llama 3.1 70B is around 45GB is size, despite being trained on likely hundreds of petabytes of data. And before you say it, they are not fancy compression algo's either, the loss is so high they would be useless.
Whether LLMs are archives of data, a compression method, or whatever else is just an unimportant technical implementation detail.
But here's another way to think about what I'm saying, in case you missed it:
Personally, I'd love to download a complete archive of JSTOR. I'd train myself, and maybe even I could even use it as input into some product I mean to launch soon. JSTOR doesn't offer a license for that, at least not to me, but I'm sure I can scrape their site or find an archive elsewhere and make it happen anyway.
Do you think I should do that? What do you think might happen if I tried?
That has nothing to do with how LLM's were trained. They were trained on countless works for which Meta, etc had acquired no legitimate right for use at all.
I will accept the argument they got the source material in a way where someone broke American law. I really do not think they've broken any laws whatsoever in terms of using it for LLM training
Isn't inducing or offering someone incentives to break laws illegal by itself? I'll admit that isn't specifically an IP law violation, but it can't possibly be kosher.
For example if a buyer of goods can reasonably be expected to know the goods were stolen, they can also be charged. Isn't this the same thing?
Nowadays the surviving public libraries might pay special prices for the right of lending books, but that was not true in the past, when they just bought the books from the market like anyone else, at the same price.
I am pretty sure that the public libraries that I frequented as a child, many decades ago, did not pay anything for a book above the price that I would have paid myself, but nonetheless at that time nobody would have thought that they do not have the right to lend the books to whomever they pleased.
Giving Meta exclusive access to those copies is the problem (which is effectively what we are doing if they are not prosecuted, or, alternatively, if we accepted that LibGen is fair use for everyone).
Whatever legal contortions used to justify this are, quite frankly, bullshit. This isn't how anything should work even if these companies can buy themselves a regulatory regime where it does.
You pay for access to materials, not using or remembering the material in its original format.
Free for me, not for thee.
Swartz was charged with 35 to 50 years, realistically faced up to 10, and was offered 6 months if he plead guilty [1]. That offer moreover wasn’t the final offer.
Put another way, it’s not clear that the law is being applied to Zuckerberg differently than it was to Swartz given the law wasn’t actually ever applied to Swartz. (Or that they wouldn’t gladly trade this lawbreaking for $1mm in fines and a negotiation over penalties where the prosecution opens with 6 months jail.)
The prosecutor acted inappropriately in that case; MIT, more wildly so. That doesn’t, however, carry over to a transgression of the law given we never got to that stage.
[1] https://www.forbes.com/sites/forbesdev/2023/02/28/increase-w...?
Has Zuckerberg actually been charged with something with equivalent potential consequences?
If not, then your statement is false on its face.
I didn’t say Zuckerberg has been subjected to what Swartz was. Swartz never wielded the nation-state level power of a billionaire—it’s difficult to imagine how he could be subjected to similar psychological stress.
I said the law isn’t being applied to Zuckerberg (or anyone who has downloaded LibGen, for that matter) differently because the law was never applied to Swartz. Given the unpopular Swartz prosecution ended Ortiz’s career, and the lack of recent criminal copyright cases, it’s unlikely anyone would attempt to apply it as they did then. To anyone, including Zuckerberg.
TL; DR If you dislike what Zuckerberg is doing, you’re probably advocating for a clarification of the law. If you like it, erm, nothing much to do here.
Merely being charged with or investigated for a crime is absolutely an application of the law.
Frankly, it's a massive boon to researchers. It's like a top-tier research university library at your fingertips, and usually more convenient than the real thing.
But the sad state of the affairs is that if Aaron Swartz does it, he ends up dead; if Meta does it, everything is fine.
Thing is, the Elsevier/Springer model makes it incredibly difficult to pay them. With single papers or book chapters in the $30-40 range, an afternoon's research can easily cost $600. (Note that the authors and reviewers don't get royalties on this, and the Editor-in-Chief of any given journal usually only makes a small stipend!)
There are services like DeepDyve, but they're intentionally gimped and difficult to use, because their user interface is 100% built around preventing you from downloading or screenshotting the papers you "rent"!
If the publishers set up a $100/month all-open-access program, and if the experience were at least halfway decent, I'd bet that a lot of people sign up. And that's not cheap!
https://hn.algolia.com/?query=libgen&type=all ("LibGen")
https://hn.algolia.com/?query=anna's%20archive&type=all ("Anna's Archive")
https://hn.algolia.com/?query=z%20library&type=all ("Z-Library")
Copyright laws are not millennia-old ethical laws that everyone agrees on (like don't steal), they are a modern human construct that were created for the greater good (incentivize creation), and we should revisit them with new tech.
How is pleasing Meta's shareholders in world's best interest.
what are you rationalising about?
Humans do that naturally (see: children)
The copyright laws are to protect profit.
Text: https://www.courtlistener.com/docket/67569326/373/kadrey-v-m...
"Meta's request is preposterous. With one possible exception, there is not a single thing in those briefs that should be sealed."
"It is clear that Meta's sealing request is not designed to protect against the disclosure of sensitive business information that competitors could use to their advantage. Rather, it is designed to avoid negative publicity."
"If Meta again submits an unreasonably broad sealing request, all materials will simply be unsealed."
"One final comment. Between this sealing request and assertions in Meta's opposition brief such as "[t]hat document expressly discusses torrents and seeding", Opp. at 7, the Court is becoming concerned that Meta and its counsel are starting to travel down a familiar road. See In re Facebook, Inc. Consumer Privacy User Profile Litigation, 655 F. Supp. 3d 899 (N.D. Cal. 2023)."
Rather famously, some elements of that administration are above the criminal code, so that's not implausible.
1- Should we develop this argument into more discussion as society and humans around the knowledge publication and the publication industry greed and the rent-seeking business model.
2- Big Corporation shouldn't just ignore the copyright law while maintaining the strongest copyright protections and going after small folks.
3- The usual argument about how LLMs training is different from people actually using pirated textbook because it is expensive (college and learning is hard and expensive specially in places like Africa).
These are different angles and I think we can try to address all of them as they are not exclusive. There are good arguments around point 3 on two sides. I don't think there is a good argument why we should allow the status quo regarding the first point though. For two, it is more complicated to even discuss specially on HN.
Will be interesting to see where this lands, because all outcomes seem to have significant secondary effects.
You're welcome! :P
I wonder who will come out on top, and whether there will be any incidental improvements for consumers, but unfortunately I can imagine an "AI training exemption" all too well.
In fact, I'd argue that Big Tech is pro copyright, because once they force the copyright holder to negotiate, the cost is irrelevant to them and they build a moat around that access.
For example, Google stole Reddit content for Gemini until Reddit was forced to the table, and now Google has a seemingly exclusive agreement around Reddit data for AI purposes.
Yep, the contradiction between them feeling entitled to use anything they want for training, while simultaneously having license terms which forbid using the output of their models to train other models is pretty glaring. Information wants to flow freely but only in one direction apparently.
I haven't been following it closely, but aren't there already court rulings saying that generative AI output by itself is not copyrightable?
I agree but for a different reason -- cost is actually relevant, in the sense that only the biggest player can afford to pay for the copyrights. If you are a small player, however your tech stack is or how good your model is, if you can't afford it, you can't compete with Google.
Now I guess it’s defended as good business and good science by so many flunkies.
Knifes edge stuff. Tech people should all be reading the books, not Mark’s steamroller.
There goes the gravy train
Their ideal outcome is that there's some narrow carveout that gives them permission to ignore copyright where they want to, while extending similar permission to as few/irrelevant others as possible.
[0]: https://www.newyorker.com/business/currency/what-ever-happen...
Indeed. But when do those intersect or diverge?
I don't blame him. What would you do? If I had a near perfect data training set of all the most useful books and a hungry AI to train, it would be the logical step.
The reason this is news is because of the stinking hypocrisy of it all. It's really the same topic as the Swartz-Altman discussion here [0], in that these giant companies want to have it both ways.
Where is Zuckerberg's shout-out for Alexandra Elbakyan? [1] Or for Brewster Kahle? Or any of the wast army of people who preserve and curate the vital culture of humanity by protecting it from intellectual property dungeons?
The colossal hypocrisy is that a company like Meta wishes to live under the protective umbrella of "Intellectual Property". It wants to stop me just stealing it's stuff and setting up a better Facebook
Were it exposed to the same rules it wishes to live by, it would be torn apart by vibrant and deserving competition within days.
All the Zuckerberg, Meta or OpenAI are doing is setting the ground for the abolition of intellectual property. They are literally the proverbial people who will buy the rope with which to hang themselves.
(Edit. that doesn't make sense insert <proverb about buying ropes that actually makes sense>)
If you train a model 20B parameters on 20T tokens, even with 1000 tokens per example, the model extracts about 1 byte of information per example. What is the value of 1 byte of copyright infringement?
The whole conflict boils down to one party having piles of money and another party having something they want. That's not an intractable problem.
Spotify began by uploading an employee's pirated MP3s, and is now valued at $92B.
There are plenty of other examples. One of the ways to success is to ignore silly legal matters, build a product people want, and worry about the legality later. It's not just AI companies, the pattern is well established.
To me, it's just "more of the same", but apparently he said the quiet part out loud, which was somehow verboten.
(Edit to add: I'm not saying "I think this is okay", but rather "this is Standard Operating Procedure for startups" -- even Reddit was seeded with fake accounts and content, to give the appearance of an active online community. This sort of hustle is a core part of SV culture, and I don't think this is going to change in a hurry.)
Excerpt:
...in the example that I gave of the TikTok competitor, and by the way, I was not arguing that you should illegally steal everybody's music. What you would do if you're a Silicon Valley entrepreneur, which hopefully all of you will be, is if it took off, then you'd hire a whole bunch of lawyers to go clean the mess up, right? But if nobody uses your product, it doesn't matter that you stole all the content.
https://finance.yahoo.com/news/ex-google-ceo-schmidt-advised...
>began
You're very obviously missing a key point here. It's rather simple: pirating is integral to "AI", as it is of the utmost importance with regards to its optimization and even to building its basic functionalities. It will never cease to happen nor is it part of some "preliminary" process in which executives "ignore silly legal matters" in order to kick-start their projects only to discard those practices once they eventually take off. Comparisons to YouTube, Spotify, etc., are invalid for this very reason.
You raise a good point. However, both Spotify and YouTube benefited from network effects and being the biggest guerrilla in the room. Can you remove the initial illegality from their later success, since the latter dependend on the prior?
What seems inevitable is that some deal is made with major rights holders, the little guy gets screwed, as has happened before.
Llama is probably one of the few LLMs that probably doesn't generate an income for Meta but I can't exactly see how other than by assisting their current ad generation.
Them being open weight isn't as good a what a "proper" open source LLM would be, but vs OpenAI which likely did the same thing it's significantly better.
On the other hand if copyright is enforced it should be enforced across the board, if I did the same thing while training an AI would I get the same treatment... Equal before the law and all that...
On the third point, I cannot legally obtain scientific paper without very significant cost to myself. My local libraries don't have a reasonable selection and even the university libraries that will let me as a member of public or even alumni still hold membership, specifically exclude scientific papers in that membership and you need to pay per paper.
Noam Chomsky, New York Times - March 8, 2023
Is it possible for an LLM of llama3/sonnet3.5/GPT4o quality to be trained on freely available works?
Are there other types of LLMs that can be trained on smaller data sets with comparable quality?
If that is not possible, and the courts shut down training on copyrighted works - what position will the "rule following" nations be in compared to nations that don't follow those rules?
While commercial motives do matter (as far as I understand), being able to find evidence and bring a case practically matters even more.
Idk, just rubs me the wrong way when there are companies making money on the exact same product sourced the exact same way right now, as Meta chooses to make it free and gets sued. Seems like we should be logical enough to conclude if they get sued, every similar company should be investigated and fined (if wrongdoing was found) as well.
Still, they don't publish their detailed training methodology, which must have immense value for them internally. Even if they choose to never exercise the "large user exemption" in their current license, they can decide to license Llama 4 under restrictive terms (or not release weights at all – better start using Meta products if you want to gain access!) whenever it's convenient for them to do so.
All of that doesn't exactly scream "public good worthy of a copyright exemption" to me, in a world where libraries, retro computing/gaming archivists and others are still continuously harassed by copyright holders.
Everyone, including them, quickly changed their tune though. Now dataset is confidential trade secrets lol
Both from a budget at FB point of view, because probably they allocate a smaller budget to training than OpenAI does, so they can't be as generous, as it's not a core part of their business. And probably also from a point of view of the publishers, they probably don't like it either: with OpenAI they can do limited time deals, but with Llama the license allows redistribution so once it's published and a lot of business activity has been established on top of it, one can't come back and renegotiate, after say 10 years.
It's in the best interest of big copyright to not have open(-ish) models in the ecosystem, they want entities they can seek rents from.
1. Because they literally said they used Books3 in the original Llama paper. The provenance of datasets used by other models is not as well documented. Books3 is known to be pirated.
2. Being free to use doesn't mitigate the authors' complaint in any way. (Compare: "I stole your bike, but then I gave it away.") The authors and artists (in the case of image models) want either a) to not have their work included in training sets or b) to be paid for that use via licensing. In either case they must enjoin the trainer of the model.
For point 1, there are employees at OpenAI who do know the provenance of the datasets used and I am sure (based on my experiences) that it includes knowingly downloaded and inserted copyrighted works. Is not one single employee of OpenAI willing to blow the whistle?
Losses for the tech companies will just mean training data gets more expensive. They're all spending tens of billions on new data centers, so there's not even a question of whether they can afford it.
Also libgen is amazing. You're missing out if you're not using it.
It's funny that most AI bros I have met were severely against piracy a few years ago.
Not that the billionaire class will ever be held to the same legal standard as the rest of us.
>FBI should go after OpenAI and Sam Altman the same way they went after Aaron Swartz.
Should now be expanded to:
FBI should go after OpenAI, Sam Altman and Meta the same way they went after Aaron Swartz.
But yea copyright is broken and is holding back humanity as a whole
So not a direct analog
It's the new radium. Expect it to be in your shaving cream without accountability and hope future generations look back on it thinking we were really dumb. We're all part of the biggest experiment in human history we just have to trust who is at the wheel.
The status quo favors large players who can navigate the legal system.
A tax on AI is stupid because the big players can dodge taxes well and have the ear of power now, so any regulation would favor them. It would only serve to prevent challengers to their dominance.
It’s the core of it all— jealousy of a creative spirit.
Artists are not machines, but living souls.
If you were remotely open enough to see for yourself, then you wouldn’t struggle with engaging in the world in a creative manner and you wouldn’t feel that jealousy but encouragement by what you see pouring out of your fellow humans as a reflection of each other.
No machine will grant you that understanding, you just have to engage directly.
It will never succeed to supplant it, no matter the billions of dollars burned to try.
Using an AI tool to create an image of a painting betrays the person who seeks to be “artist” by short circuiting the practice that leads the prospect to their path of enlightenment through mastery.
In our constricted 3d world there is no circumstance where an algorithmically generated image of a painting will equally serve the prospective artist in its procedural work on the prospective artist, internally. There is no other pursuit in art, and any pursuant will come to that conclusion in any number of ways but always through submission to the course of mastery (for which there is no shortcut).
Worse, the companies at the helm of this side of the technology are pushing it in order to stand middle man to humankind’s modus operandi-to create.
Keep your mind open to perspectives beyond the software industry.
Oh, you said zero? How beguiling. Maybe there is a difference.
BTW, plenty of humans have memorized copyrighted material, such as song lyrics. Do you think that should be prohibited? Maybe the difference isn’t as great as you think.
OAI can put a dumb IP filter on ChatGPT output and resolve the case. Training plays no part here.
That regurgitation is merely evidence of those two, and so putting a filter on the output explicitly does not resolve the case.
LLM's are not data archives, I don't know how many times this has to be repeated.
> Defendants’ generative artificial intelligence (“GenAI”) tools rely on large-language models (“LLMs”) that were built by copying and using millions of The Times’s copyrighted news articles, in-depth investigations, opinion pieces, reviews, how-to guides, and more.
https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
This case is still ongoing a year later.
NYT isn't suing because LLMs will print out NYT articles for free. They are suing because LLMs are poised to be better/more favorable news reporters than them. It's a long term survival case, not a copyright one (despite that being the weapon used in the fight)