It sure looks like Meta stole a lot of books to build its AI
lithub.com
lithub.com
Wow if that's the opener, I expect the rest to be SUPER emotionally charged
The next paragraph...
https://theintercept.com/2025/01/09/facebook-instagram-meta-...
Haven’t tried it yet
In other words, unless you believe Facebook is part of State or acts on behalf of the State, “you don’t get to shut down hate speech” is mereley your sentiment, not something Constitution requires from them.
If Mark Z actually called hate speech “protected speech”, it is something he is free to do, but Facebook is also completely free to suppress hate speech.
* https://www.nyclu.org/commentary/column-applying-constitutio...
"Billionaire oligarchs" only became a problem when a few of them became neutral or moderately right-wing.
It's the "anything can be a slur is you say it with enough hate" thing. You're just giving a name to the group targeted by it. The content isn't the thing that matters, it's the dehumanization and harassment.
From adults, yes
This should get to the heart of it: what biological reality are the leftists discussing trans issues denying?
I’ll also respond.
> There are indeed people and whole groups that will call you transphobic if you do suggest certain things about sexual dimorphism between men and women
I expect such people would call you misogynist, not transphobic. I also think it’s mostly down to delivery. People who have issues with trans people often talk about these things in certain ways, so people assume anyone who talks in such a way is a transphobe.
> if you at all question the idea of often very suggestible adolescents being easily allowed to go through the process of gender reassignment
It’s not easy to my knowledge. Adolescents are never given sex change operations, those aren’t even typically given to minors. The interventions are limited to puberty blockers, which are highly reversible, hormones, which they receive after years of therapy to confirm it’s not just a phase, suggestion, etc. and which can still be largely reversed, and social transition, (dressing/presenting as the opposite sex) which hopefully anyone would be fine with. Which part of this process do you find contentious and why? Again, I’m happy to discuss this, and there’s nothing wrong with asking questions about it. In fact, I think it’s extremely important to ask questions about this because it may help to protect children. The issue is that I mostly see people bring this up not because they know something about this process and dislike it, but because they don’t think people should be trans.
In case of Facebook, it was originally designed literally to compare the attractiveness of female students; but no, it’s not even remotely as bad, and comparing it to the Holocaust trivializes an incident where masses of people were murdered in an industrial fashion.
Possibly the point the author is making is that Zuckerberg never had an original idea or one that could be the subject of a business plan. He copied the "Hot or Not" websites that had come before. Further, the photos of students used for his "Facemash" website, i.e., the "content", were downloaded from the university's computers, not uploaded to Zuckerberg's computer. Initially, he downloaded and used students' photos without permission.
https://web.archive.org/web/20250115010420if_/https://www.th...
This pattern continued when he copied the idea of an online "face book" for the university which the university was already working on; adopting the name "thefacebook.com".
From the document production in Kadrey v Meta, it appears the pattern of copying still continues. Meta is still downloading and using others' work. Initially, without permission.
Comparing copyright infringment to assisting genocide is absurd. Although Meta may have assisted in ethnic cleansing
https://www.amnesty.org/en/latest/news/2023/08/myanmar-time-...
there is no reference to it in the article. Perhaps because, unlike the story of "Facemash", it bears no relation to the subject matter: copyright infringment.
Walden was first published in 1854. At the time, the maximum length of copyright in the US was 28 years (14 at first + 14 on renewal).
Notions of "fair use" in the US can be traced back to the mid-1800s, too. There were court rulings, but fair use was not codified into law until 1976. Non-profit educational use was explicitly called out in 1976 also.
Photocopiers were first patented in 1937.
[1]: All your favorite authors, journalists, anything indie.
Could you clearly speak your point?
Those people who conflate them deserve it. You and me don't.
> there's really not much you, I or anyone else can do
We can make our own community. And Bluesky is very much not it.
I agree with you until this part. There comes a time where I don't think I deserve to get my eyes poked out just because other people find that fashionable.
Extrapolated out into some new future a hundred years from now when we have embodied AI humanoids walking alongside us, would it be weird if those humanoids were barred from buying a new book or charged a different rate than the humans they coexist with?
I’m still deciding how I feel about some of this too.
I'm not even against this to a point. The issue is what comes after. The monetization. The enshitification. The derivatives in place of real creativity.
The only way to prevent the things you are worried about is to let anyone train a model, and then compete to make the best product.
The enshittifying monopolies and big copyright holders are the only ones that would benefit from locking down training data or regulating AI compute.
They're not being charged, that would be a vast improvement over reality.
But if they had done that, I bet they would have been sued anyway.
“Because you wouldn’t have sold it to me. Or even if you would have, you would have put such onerous terms on it that I’d rather take this path.”
(Not saying this makes it right)
So just buying copies wouldn’t have helped them.
For a more direct counterexample, I can memorize something and type it back out, but if it is copyrighted the law doesn’t make an exception just because it passed through my head.
I do agree that we should encourage human creativity. But if AI isn't making copies, and the output of AI isn't awarded copyright (as is currently the case) then I think humans still have sufficient reward.
There will be a lot to figure out over the coming years.
On the other hand, they torrented books and then open sourced LLM weights. No punishment is too severe for that!
If you still don’t understand, I strongly suggest watching Max Headroom, “Lessons”, which you can get here:
Now, would that be a fair use of the books?
Ok, it doesn't tell much about AI and fair use, but I find it funny that your thought experiment is actually something you can do in real life.
There's no reason for Harry Potter for example being 10000 times more valuable than a book on quantum mechanics only because the former is more popular and the latter is on a more obscure topic.
You have bought the text so you have the readright, but you do not the copyright.
You do however, have the right to make derivative works based on the contents of the book. You reading a physics textbook doesn't mean you can't write a blog post about gravity or whatever, and you reading harry potter doesn't mean you can't write a series of fantasy books involving a young wizard trying to fight an evil wizard.
> The application was denied because, based on the applicant’s representations in the application, the examiner found that the work contained no human authorship. After a series of administrative appeals, the Office’s Review Board issued a final determination affirming that the work could not be registered because it was made “without any creative contribution from a human actor.”
That just means whatever they produce can't be copyrighted, not that they can't produce derivative works. Courts have upheld the right for google to produce thumbnails of copyrighted works, even though the procedure for producing thumbnails is done by a computer and thus can't be copyrighted.
Maybe we will end up agreeing that we just want to stick with those same laws for machine consumption and creativity. But maybe we won't since they are quite different things.
If a software service had legal protections like that, sure, I could build one that returns you any book you request and say that the service had integrated it into its worldview. Who can check, eh?
* Actually, in some countries you could be in trouble for reading a book and incorporating it into your worldview, to say nothing about quoting it, but let’s set that aside.
Not a relevant factor when it comes to copyright law. Fair use (the law that's most applicable here) applies regardless if you're a student using incorporating news articles into your work, or google making thumbnails and displaying them on their search results.
Furthermore:
> Examples of fair use in United States copyright law include commentary, search engines, criticism, parody, news reporting, research, and scholarship.
I do not see “automated generation of derivative works of arbitrary nature” in it.
The “automated” isn’t really key. If you read a book, and learn from it, and are able to use that knowledge in other contexts, should you pay a licensing fee? It doesn’t matter if “you” is a human or machine.
The point isn't that AI training is legal because it's like generating thumbnails. That is being argued in the courts right now. The point is that fair use exemptions isn't limited to "being a conscious human being enjoying human rights", as google generating thumnails and snippets using computers shows.
https://en.wikipedia.org/wiki/Perfect_10,_Inc._v._Amazon.com....
> Examples of fair use in United States copyright law include commentary, search engines, criticism, parody, news reporting, research, and scholarship.
Those are examples, not an exhaustive list. It's not even something that Judges are supposed to compare against when deciding whether something is fair use or not, see: https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors
Sure. However, my point is that this is not fair use*, so other principles need to be applied. Whether legal systems in various countries find that fair use applies here or not, I agree we are yet to see.
* At least in cases where it’s an LLM operated at scale for profit (which I suppose would not hold for Meta’s models if they were truly open, but that’s not the case if they require obtaining a license in some conditions).
This isn't a complete argument. Most of AI companies' argument relies on the fact that AI models are "transformative". That's a plausible claim, and as Perfect 10 v. Google, and Authors Guild, Inc. v. Google, Inc. has shown, being a for-profit company is hardly a disqualification from getting fair protection.
But sure, the “transformative” argument is the one that could apply (and even I believe Google used it to argue its case), if it can be shown that an LLM can not verbatim reproduce a given work (which, incidentally, is something that you, a warm-blooded fleshy human with agency who has the freedom to read books, cannot do, but LLMs were shown to do).
That said, relevant laws existed before LLMs, and may are outdated. If the goal is to balance reasonable uses while protecting original output of authors that ultimately drives innovation and creativity, I am not sure if the preexisting laws are continuing to fulfil their function, but that’s my opinion.
You have to try pretty hard to get LLMs to reproduce a work verbatim, especially any lengthy passages that aren't famous (and thus re-quoted on the internet a bazillion times). Moreover just because LLMs can reproduce a work verbatim if you try hard enough doesn't mean it's not transformative. Google search snippets and google book search has been ruled "transformative" by the courts, but if you tried hard enough you can use them to extract the entire work.
>That said, relevant laws existed before LLMs, and may are outdater. If the goal is to balance reasonable uses while protecting original output of authors that ultimately drives innovation and creativity, I am not sure if the preexisting laws are continuing to fulfil their function, but that’s my opinion.
AFAIK the era of mining the public internet or published works for AI training data is over, or at least coming to an end. Everything that could be mined, has already been mined, and besides, the internet is getting increasingly polluted by AI output. Private training data is where it's at now, whether it's sourcing document troves from companies (eg. emails, documentation, source code, etc.), or paying "AI annotators" to produce training data for you. If the argument is that human authors should get a cut of AI profits because their works were "stolen" to train the models, this is going to be a increasingly losing argument, because it doesn't have a leg to stand on for private training data.
The argument can be made that LLMs could not be created without expropriating the original works of all the authors they were trained on, and that argument would in fact be true and have quite sturdy legs as far as I’m concerned.
It’s not a historical instance of forgotten times, it started less than half a decade ago and I would be surprised if it’s not still ongoing (your argument about synthetic training data is forward-looking).
That makes as much sense as "American industry was built on the backs of British inventors (back it the day it was the "China" when it came to IP), so Britain should get perpetual (?) royalties from the US economy".
You are arguing that doing something that is legal if being done by humans is not ok if it is done on computers running an LLM.
I see no difference with cryptotokens here, the human has freedoms to do things and the human is responsible for them if those things are bad. (Just unlike LLMs, theft of property and all that is kinda always a crime, unlike reading a book in a shop without buying.)
So I expect to see that either you are no longer allowed to own computer software
Or a return of slavery.
Also if we find indecent portrayal of minors in a data centre I expect that we treat it as a strict liability crime and the entire data centre or corporation that owns it gets a long prison sentence, just like a human would. However that is suppose to work.
That... doesn't make it okay...
> A lot of them I didn't even pay for, I borrowed them from libraries or friends.
This 2nd sentence doesn't fit your first. What is your message?
But I'll try to articulate it anyway. The people who created the data all these models trained on, be they artists, writers, or even programmers, created a lot of, if not most of, the value that is now being derived from these models. Instead of being rewarded for their part, a lot of folks here seem very content with casting those people aside and letting huge corporations take everything, while building a system that is trying to make people creating actual things that have value have a much harder time surviving off their trade.
It's very gross to me that people are defending Meta here, and seem to be okay with capital eating all forms of cultural expression while giving nothing back.
That said, my personal belief is that if the books weren’t legally freely available all that Meta owes would be how much the book costs. Each individual book would be such a small part of the model that it’s barely distinguishable. I’m sure image models have been trained on some of my professionally taken photographs and I don’t care one bit.
I’d argue that’s the cost of the books was how much it was worth before AI models, and the authors themselves didn’t create the technology. Therefore the added value of the technology has absolutely nothing to do with them. If book publishers/authors decide to have different pricing in the future to take the tech into account that’s their right.
Why isn't stealing and knowing you're stealing penalized more than the cost of the item? If the world worked this way everyone would steal.
> Each individual book would be such a small part of the model that it’s barely distinguishable.
Needs to be proven (and also impossible to prove how much of the value of the model comes from the classified material)
Knowingly obtaining millions of copyrighted materials that were posted illegally simply to serve your own financial interests, might very well qualify.
The problem here is the tech industry sits on throne of riches built by IP law, so it doesn't sit well when suddenly it's "good for thee but not for me." If we're going to cherry pick, how about we walk back to a view that software isn't copyrightable and the copyright term is 34 years?
Maybe if the copyright system wasn't so extreme, people would have a more balanced view of the system and show more support.
I don't think anybody really believe Meta is doing this as a charity.
For example, if Facebook employees broke into the home of a renowned author and stole private copyrighted materials they then used for training their LLMs, should the Court, in analyzing the Fair Use factors, disregard the illegal nature of how the copyrighted materials were obtained?
I believe it unlikely a Court would be willing to reward such behavior.
Once that principle is resolved, the next step would be for the court to consider whether it would make any difference if Facebook employees did not engage in the direct theft, but acquired copies of the stolen materials from the thief with full knowledge they were stoken.
If the court believes both #1 and #2 would be unacceptable, their analysis would then proceed to consider if there were differences favoring Facebook if Facebook acquired millions of copyrighted materials via a notorious website widely accused of illegally posting unauthorized access to copyrighted materials.
I suspect this will likely be the central issue of the legal debate. And, I for one, do not think Facebook has a very strong legal argument. Going back to the first step of the anslysis, I would be shocked if SCOTUS would be willing to state that how the copyrighted materials were obtained is irrelevant to the Fair Use analysis.
cough https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
I thought they give it away for free?
Meta isn't a charity - even if they're not profiting from Llama today, they believe they will at some point.
(Not trying to select + copy, just trying to indicate how far I am through the document)
Isn't it just a matter of time before we see content creators weaponize their sites? They may present text as images or require a user to solve a pictogram cipher etc.
Effectively turning the web scraper into an attack vector.
The CSV score would be weird for that.
Did you mean CVSS?
I think I ate to many Christmas food....or maby it was the chocolate bars and Hersey kisses I consumed.
With uBlock, though, I have 20+ blocked items. What's the point of having so many scripts that do absolutely nothing visible on the page?
Is there any plan to recover the books and return them to their owners?
Maybe the Library system does apply here though, that all AI trainers need to buy 1 copy of their book so "their child" can "learn"?
Why can't they borrow the book from the library?
What you think isn't law, it will be decided now.
> Should Meta do the same?
Well, they didn't do that. They stole the books and knew they were stealing them.
People read books to learn and then use that knowledge with no contribution or even recognition of the publishers or authors of the books. Sometimes people even quote books without payment to the authors. Sometimes the books are used or borrowed and the publishers and authors don’t get paid. How is it different when an AI does the same thing? Or when a human uses an AI to do the same thing?
Even if Facebook is in the wrong here what is the remedy? Would it be ok if they purchased used books and scanned them?
More discussion: https://news.ycombinator.com/item?id=42651007 https://news.ycombinator.com/item?id=42673628
Zuckerberg sucking up to the right wing because the cultural landscape is swinging right is also not surprising.
That's the gist of your argument. Take a moment to consider what can be justified by that same reasoning. Just about anything.
People work hard on what they write. Many writers struggle to generate a reliable income. Yet to you, it's okay for a competitor to take that work without compensating the author, then use it in a product they will sell. Because "information wants to be free".
It breaks my fucking heart how little our society cares about the compensation of artists, authors, and creative people generally; all while gobbling it up and expecting it for free.
Why is software different? Where should we draw the line between fair use and copyright violation when it comes to AI?
No, you want it to be free. The people whose labor created that information didn't want it to be free, especially for some gigacorp to launder while hiding behind fair use.
I have a copyright on information I have created, but would never enforce it because I want it to be free and consumed by both humans and AI
I'm also not sure what you mean by "laundering", though it seems like there is the prevailing belief among a subset of artists that the AIs are regurgitating their works (which would be a copyright violation) rather than just incorporating them into their world representation (learning, fair use). While there have been instances of the models outputting certain texts, we have not seen a new story about this in awhile. There are multiple remediations for this issue. I don't find banning copyrighted material from learning material a good one. It's like book banning in schools if you ask me
Weird cope about AI, not a very smart one.
1) Every author I've seen discuss this technology completely despises it. Using it to write your book is a reputational risk.
2) Putting the words on the paper isn't the hard part. I can imagine many authors wouldn't use this tech even if it was morally and socially okay because they care about whether their voice is getting through in their work.
If I train cp on movies released this year and then release each model open source, is this now legal?