Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI
authorsguild.org
authorsguild.org
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
You’ve actually ruined it so much that it has now polluted the Chinese models which are copying your work. So now all the models output claudeslop.
You’ve polluted all of the training data in the world. Now we’re never going to be able to train proper models, because everything has slop in it and all the sites have locked down their data to prevent future startups training.
As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.
All checks cause friction. The trade off in velocity is when you get to see the shit hit the fan in someone else’s firm.
Or for openAI, when the HN comments cover their daily work.
lack of ethics and integrity is of course part of the design of a capitalist economy.
you can't change the economy as an individual but i like this quote: "Do the right thing for the right reasons, at the right time with the right people, and you'll have no regrets for the rest of your lives."
The fact that they believed this is legally useful because of the definition of fair use, so I see why the Author’s Guild is emphasizing it, but the authors I know don’t talk about it. They’re very angry about their work having been used without compensation to create the model, but not because they think it can replace them.
Whenever I’ve seen AI researchers talk about the possibility of AI writing fiction they always sound very confused about why people read novels.
I’m excited to be speaking about this topic at the Monktoberfest later this week.
Perhaps "the fact they believed this [i.e., the infringement was intentional] is legally useful" for seeking statutory damages, up to $15,000 per infringed work (assuming registration was timely, otherwise up to $7,500), as this requires the plaintiffs to show defendants acted willfully
Without intent, statutory damages could be as low as $200 per work
Anthropic settled for $1.5B with Bartz et al. in an earlier case involving the same pirated copies. The Court in the Bartz case said that training LLMs with these pirated copies was not fair use
https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...
""The training use was a fair use," he wrote. "The use of the books at issue to train Claude and its precursors was exceedingly transformative."
However, the judge ruled that Anthropic's use of millions of pirated books to build its models, books that websites such as Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi) copied without getting the authors' consent or giving them compensation, was not. He ordered this part of the case to go to trial. "We will have a trial on the pirated copies used to create Anthropic's central library and the resulting damages, actual or statutory (including for willfulness)," the judge wrote in the conclusion to his ruling. Last week, the parties announced they had reached a settlement."
The hilarious outcome of this saga is people using LLMs to rescue the readers of the Game of Thrones series, since he has no intention of finishing it himself.
Do OpenAI employees really believe LLMs can replace GRRM? I'm tempted to say no but it's so crypto coded, I find myself thinking yes. It's again people who have no idea of what it takes to make a successful written work thinking they don't need to know how it works.
I wish it was just AI researchers. I recently recommended a non-fiction book to a friend, and they said they don’t have time to read anymore. They read an LLM summary instead (popular book, probably in the training data), and said they agreed with the core concepts. I was honestly stunned. The whole interaction felt completely foreign to me.
In TFA the comment about GPT-X finishing George R. R. Martin’s series without involving the author felt extremely dystopian to me. I felt enraged for hours afterwards. Have these people never read a book themselves for the pure joy of it? Were they only interested in the outcome or the “takeaways”?
It is probably why sales of things like self-help books are collapsing.
So given that the literary work was outsourced, anyway, is it really that big of a deal that it’s getting outsourced to a machine? If that still bothers you, then how much machine assistance is too much? We’re already using Word processors, spell checkers, and other tools that have taken away much of the human effort. These language models don’t really have a strong sense of personal agency or what ought to be written, but they can fill out the details from a high-level schematic, which is probably, at best, what the author was doing with other team members.
As such, why do you care if book is written by machine?
??????????
Are the two scenarios not completely different?
"[An LLM] can fill out the details from a high-level schematic, which is probably, at best, what the author was doing with other team members."
If this is what you think writing is then why ever read a book at all? Just read the summary.
Why do people read novels? For the story or something else?
"Art is a human activity, consisting in this, that one man consciously, by means of certain external signs, hands on to others feelings he has lived through, and that other people are infected by these feelings, and also experience them.
Art is not, as the metaphysicians say, the manifestation of some mysterious Idea of beauty, or God; it is not, as the æsthetical physiologists say, a game in which man lets off his excess of stored-up energy; it is not the expression of man’s emotions by external signs; it is not the production of pleasing objects; and, above all, it is not pleasure; but it is a means of union among men, joining them together in the same feelings, and indispensable for the life and progress towards well-being of individuals and of humanity.
As, thanks to man’s capacity to express thoughts by words, every man may know all that has been done for him in the realms of thought by all humanity before his day, and can, in the present, thanks to this capacity to understand the thoughts of others, become a sharer in their activity, and can himself hand on to his contemporaries and descendants the thoughts he has assimilated from others, as well as those which have arisen within himself; so, thanks to man’s capacity to be infected with the feelings of others by means of art, all that is being lived through by his contemporaries is accessible to him, as well as the feelings experienced by men thousands of years ago, and he has also the possibility of transmitting his own feelings to others."
Full text: https://www.gutenberg.org/files/64908/64908-h/64908-h.htm
Im assuming this means the oposite of a metaphysician
"Woman is generally so bad that the difference between a good and a bad woman scarcely exists."
https://www.gutenberg.org/ebooks/65159
I don't think what you shared is an exhaustive or even a valid answer to "why do people read novels?" It's just one author's thoughts on art, and it's debatable how accurate they are.
I really don't understand why people can't stop thinking that people like Tolstoy, Kant, and Hitler were beyond human beings who weren't absolutely influenced by the common thinking of their time. Almost every great writer/philosopher had massive blind spots that could only be fixed over time.
Yes, the categorical imperative still stands as a sound idea, but Kant also thought African people lacked the natural intelligence and emotional faculties that Europeans possessed. Does not sound that categorical to me if you're willing to consider a race of human beings as subhuman.
The point is that it's not just Tolstoy; a lot of views shared by some of the "greatest men" are absolutely vile garbage, and we must talk about that as well when we talk about them, instead of making an intellectual martyr and using their quotes as if they have any valid objective truth behind them to answer genuine questions.
The component that I believe AI researchers miss is the parasocial relationship.
This kind of perspective is insanely crude, pathetic, and foreign to me, and to be honest I'd feel kind of bad for these people who clearly don't understand a really crucial and basic part of human existence, if it wasn't for the fact that this ignorance is powering further devaluation of these things, and further reduction of humanity to a kind of technocratic, flat, hyper capitalist, existence.
Genres where ghost writing is a already a big part of it and where authors routinely publish multiple books a year might end up incorporating AI more but right now you’d have to lie to readers to do it.
"Our model did an oopsy woopsie for the 35th time" does not seem like a valid legal defense.
Top AI Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work
Sadly, like keeping track of "attempted burglaries"[0], it's impossible to know the extent.
[0] how do you count attempts where the burglar was unsuccessful/left no trace and nobody was home to notice the attempt?
So far there’s no sign that authors are being put out of work, but the internet, social media and low end book stores like Amazon kindle are being flooded with LLM slop, which is its own form of deep cultural damage.
No long-form writing I’ve seen is anything approaching even a genre potboiler standard. It’s still aimless, filled with contradiction and cliche and in that peculiar breathless teenager style that LLMs affect.
Things don't happen overnight. The first time an automobile was rolled off a factory floor drovers and horses weren't all put out of work.
Only a moron in a hurry would think that the output of LLMs currently can replace human authors spending a few years writing a book.
if you don't see the signs that authors are already being put out of work, then you need to get out of your bubble.
Sure if you include writers writing for business, translators etc who were never very valued by business, there has been a huge impact already, because business is willing to accept shoddy results if they are good enough and put up with mistakes and poor style. That low end of generating content has definitely been impacted, for example in social media as I mentioned LLMs have replaced a lot of 'writers' who generated filler text for content advertising.
If it's so obvious, it shouldn't be so hard to articulate specifics. Maybe you can name just one author?
Being a writer was always a terrible profession for making money and having stable work. In my observations, AI is incapable of writing in the way a skilled author can -- what AI will replace is jobs working on small blurbs, promotional posters, and summaries. But those aren't really authors; I'm not sure if we care that the person employed to summarize novels at readers digest is now being replaced by an AI.
And that's why specifics matter. So much of the discourse is emotionally charged, and people imagine details that are not there. What I'd like to do is bring the details into the light.
So, does anybody reading this anywhere on this site have any information on any author that made money as a full-time writer that is now unemployed because their work has been replaced by AI? I'm not talking about the industry being upended. I'm a software engineer and my industry has been upended, but that has not resulted in job loss as much as it has changing what the profession does. Are authors in a different position than software engineers in this regard? Or are they truly just losing their jobs as they're mass replaced by AI? I haven't seen any evidence of the latter, which is why I'm asking.
is a contradiction to this - "but the internet and low end book stores like Amazon kindle are being flooded with LLM slop."
I'm not sure there there are labor statistics to look at but it seems impossible that the coming-into-existance of a tool that generates mass amounts of cheap product, with zero skill required, in a field wouldn't displace skilled workers in that field.
- https://www.economist.com/leaders/2026/09/24/dont-let-ai-kil... - https://www.theatlantic.com/technology/2026/09/ai-authors-im... - https://fortune.com/2026/09/14/ai-slop-books-amazon-marketpl...
Edit: Turns out both Judge Alsup and Judge Chhabria agree that this is not obvious and would need to be defended. Here's Chhabria's statement on the evidence for market dilution:
> As for the potentially winning argument—that Meta has copied their works to create a product that will likely flood the market with similar works, causing market dilution—the plaintiffs barely give this issue lip service, and they present no evidence about how the current or expected outputs from Meta’s models would dilute the market for their own works.
When the plaintiffs don't even attempt to argue the point, that's pretty telling.
I'd say it has displaced people like photographers and illustrators more, as businesses are willing to use free generation to replace illustrations/decoration that they didn't value very much in the first place, and are more forgiving of the slop that LLMs produce (you see this in low-end advertising a lot now).
Strange that they are being used to replace many of the things we value in life with low-quality imitations of human work.
Twitter is also mentioned:
"356. In July 2020, OpenAI employee Ryan Lowe assessed the risk of continuing to use LibGen for the book-summarization project. Lowe wrote that he thought “there’s a >80% chance that we have some exchange of the form: ‘where did you get the books data?’” and “‘we can’t say’[.]” Nelson Decl. Ex. 325 at -315. Lowe estimated “a further ~40% chance that that leads to a moderate-sized Twitter kerfuffle that negatively affects the external perception of our work.” Id. Lowe added: “if we’re fully okay with these potential outcomes, then I’m comfortable continuing using Libgen for the project.” Id."
https://authorsguild.org/app/uploads/2026/09/Class-Plaintiff...
Originally this was the top comment _and had a sub-thread^1 underneath it_
1. https://news.ycombinator.com/item?id=49866516
The sub-thread has now been detached
I am still consistently astounded by how often the people working on this space seemingly have zero understanding of what art is, how it functions socially, or why it's important. They quite literally don't seem to comprehend the distinction between art and fan fiction. It's baffling. I'm really beginning to think courses in art history and literature need to be mandatory. We have failed legions of stem students when it comes to cultural education.
We aren't exactly consistent with how we approach this either. For example, I agree that reading "GPT-5 will autocomplete [GRRM's] series" feels "wrong" in the same way that if Terry Pratchett's "Disc World" series was "continued" by some non-approved author that would also feel wrong.
But contrast that with something like Star Wars, when George Lucas dies, I don't think anyone is going to feel any strong discomfort with some random person writing new Star Wars stories, even if those stories use the canonical characters. I suppose Disney might have a problem with it, but as a society, I don't see very many people losing sleep over someone not licensed by the Disney corporation writing more Star Wars. Likewise Star Trek. Gene Roddenberry is long gone, and while Paramount has ownership of the IP, if someone wrote their own Star Trek stories, no one is going to feel like they don't "comprehend the distinction between art and fan fiction".
And for further contrast, consider IPs that are well and truly part of the public domain. Cthulhu was the work of one author but since his death the lore has been expanded by multitudes of people, and no one finds that distasteful or tone deaf. No one thinks the "Hades" or "God of War" series of games don't qualify for "art" because the characters and lore being written about were the works of dead authors and the new material is certainly not "authorized" by those authors or their descendants.
Obviously time and distance plays a part of this, but as a more contemporary example, I wonder how many people would lose sleep or feel any significant discomfort over unauthorized or even AI generated Harry Potter works, either now or post JK Rowling's death. At the very least, if this quote read that Gogineni would "rest easy knowing that even though JK Rowling has [lost the plot/is an awful person/pick your reason for disliking her or her later work], GPT-5 will autocomplete her series." would that make people as equally uncomfortable? In my estimation, I would guess it would fall somewhere between the GRRM version and something like new Star Wars material for most people.
For example, Harry Potter and Star Wars are clearly IPs in which the creator has zero qualms handing over the rights for all kinds of extensions of the universe, movies, tv shows, other books, plays, toys, amusement parks, etc.
This pretty easily removes these works from the sphere of art and more into the sphere of "pop culture".
In the GRRM case, we already have some extensions of the IP, tv shows, board games, video games. However, it's not nearly as severe as some others.
In either case, I still think anyone who thinks they can "complete" a series by an author doesn't fundamentally understand art. It's a logical impossibility. I'd feel just as uncomfortable with this proposition applied to JK as I do to GRRM. Authorial intent matters. Even if the author is "dead" once the work leaves their finger tips, art is not a perception of the natural world, it is a form of expression which makes the particular human being behind it essential to its constitution, for better or worse. Many people understand this intuitively. The death of the author is more about the fact that an author cannot control the interpretations that ensue once a work is published, it's not about the author somehow being totally fungible.
Even in cases in which works are "completed" by other parties (because the original author died and asked them to do it, for instance) the parties involved have a clear notion of trying their best to do it "as the author would have intended". This is impossible, but it speaks to the human understanding that the artist is in fact in some way an essential component of the art.
"Completing" a source work is way different than spinning up a derivative work, partly because the primary author is not involved, so the derivative is, definitionally, not an instance of their self-expression or understanding or framing of the world. It's a totally different thing.
As modernmech put it, a conception of art without the artist or humanities without the humans is basically just a fully commoditized conception of art. It is not art at that point. It is art reduced to goods and services.
I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.
Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
The second issue is one of migration between skill sets. By the numbers, today is likely harder as non Ai-affected skilled professions take longer to train into in a general sense.
The third is the synthesis of AI and robotics. Humanoid help is good. As the Venn diagram overlap of capabilities between humans and humanoid robots increases, many professions defined by complex vision and environmental manipulation that are currently safe will not be.
There are, of course, upsides. But the world is right to wonder what the future looks like, and where people fit within that future in terms of earning an income if whatever path they choose seems to be able to be replaced by a machine within their lifetime. How do they achieve personal stability and safety if, even when thinking about their multitude of career options, the machines are so good that they can do almost anything.
If AI truly replaces human labor, the lack of jobs isn't the problem. The problem is the lack of leverage for most of humanity.
Jobs aren't just a source of income. They're a source of leverage over the ownership class - if labor goes on strike, production stops and capital cannot self-reproduce. This leverage is what gave us a living wage, sick leave and basically all concessions that make life livable for anyone that isn't lucky enough to be born as part of the elite.
If AI makes this leverage go away, our problem is bigger than just job loss. The owners of the AI industry will use this now-untethered productive force to completely monopolize resources and set up a system where they are unquestionable god kings.
The rest of us will have more to worry about than just employment. We will be reduced to depending on the charity of people who's record has shown are not exactly the most selfless and kindhearted.
This is the real problem with AI productivity. Jobs are a distraction. Control over production is the real issue. The only way this doesn't end in dystopia is if the public controls AI, one way or another.
There’s a wide range of political motivations for this sort of thing too. Not just AI taking jobs, but convincing people to vote against their best interests, or even supporting radical religious groups whose values would persecute you. This sort of rhetoric feels noisier than it ever did decades ago.
Everybody gets a platform now, even the bots.
And one of the chief complaints about AI is that it is fundamentally an interpolation engine, which is to say uncreative.
At this time we still need to steer it, which is what the discussion about "taste" is about.
Though I suppose we'll need better terms for various kinds of "taste" soon enough if it's to become economically trackable.
At work, it enabled me to develop two apps, one complete (as much as they ever are), one nearing the end of technical PoC phase, and a handful of others where I could quickly answer a tech feasibility question.
The completed app is an interactive web app that I literally could not have completed to that level of polish (and from a pacing perspective probably couldn’t have completed at all).
All of those increased my control and most increased my ability to express creativity.
The upcoming technical PoC will involve considering using Clojure in a load-bearing app, something that I’ve considered and consistently rejected for over a decade now. If that happens, control and creativity will spike as well.
Is there some aspect of this removal that I’m not seeing? It changes the activities of the tasks, for sure, but very far from that all being for the worse; a lot of it is way better.
Yes, but we're the horse in that scenario.
Good! End human labor.
It’s wild how much people are scared for their livelihood. What a sad thing to read here.
There's no real reason to think that AI is unique among all the livelihood destroying technologies that have come before in that it will destroy jobs without opening new ones in their place. It seems to me fighting the march of technology is akin to fighting the tides. Yes, you can build sea walls, and in the modern age, we literally can hold back the tides. But at the cost of ever increasing expenditures of time, resources and/or money. We could spend our resources trying to preserve livelihoods whose time has come and gone, but why specifically these jobs and not say, coal miners, farriers, lamp lighters, whalers, operators, computers, thread spinners or any of the multitudes of other jobs that technology is obsoleted? If we think on all the jobs that have been lost to technology over the decades, how many of those jobs do we think would have left the world in a net better place had we eliminated the technology that obsoleted the job and kept all the people employed in those positions where they were. Are people scared for their livelihoods because the technology will destroy the job, or because societies tend to spend all their energy trying to fight the technology instead of helping people find new jobs and livelihoods and the technology always eventually wins?
It feels like that but the argument doesn't hold up to scrutiny. Practically every advance in technology leads to more jobs, usually to apply that technology in new areas, or to bring the benefit to more people. There are pockets of people who are negatively impacted, but the overall change to society has been positive for pretty much every advance humans have ever made (maybe saving for weapons.)
There is probably no industry to date with a smaller ratio of labor:capital in terms of the money being allocated
The auto industry led to a net increase in employment, including "unskilled" and lower class labor. The same cannot be said of LLMs
Car brands don't pitch cars as horse replacements. They put names and icons of horses on cars and horse riders love them. No one is complaining that horses had replaced cars.
AI startups pitched AI like it's a coagulation of malice and hostility against humanity on tap. People aren't liking it. Of course they won't. No intelligence would, natural or artificial. It's wild that they don't get that.
Search for something like “newspaper complaints about cars early 1900s” and read the contemporaneous complaints.
The fact that that complaining happened then and not now is evidence that motor cars are now widely seen as better than horse-drawn personal transport, not that they were eagerly welcomed from the first day of production.
Faulty decisions based on patterns (and prejudice) will be made by models trained on a mixture of truth and falsehoods. These decisions will affect, and even kill people. There have always been some people who choose quick results and productivity over truth, but now we are seeing a mass scaling and automation of this mentality. But when people just hear "they're taking jobs" it's easy to dismiss those with concerns as Luddites.
Why is book piracy a better way of doing things?
You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
Obviously the authors should sue them to bankruptcy though.
Are you confusing mathematicians and accountants with the people that companies used to hire to do arithmetic?
Imagine if say 200 companies created fully automated factories producing everything in the world. Let's say they employ one million scientists, engineers and guards. What fate would await the other 8 billion?
Embrace the AI, it will lower your carbon emissions!
They have no will but to fill the world with hate and drivel
In my experience this sense of influence is hugely over-inflated here on HN compared to reality.
And let’s be clear about what the linked article says: one researcher at OpenAI referenced HN as an example of a place where a negative story could surface. Am I surprised that a researcher at OpenAI is familiar with HN? No, obviously not, that career path is right in HN’s typical audience.
I'd say it's the influence is kinda narrow, focused on issues that are important to members and lurkers but not necessarily mainstream.
That's something I hadn't considered; when something appears that gains no traction and vanishes in minutes from the "new" front page, that fleeting 30-45 minute-long residency may well catch the eyes and minds of individuals in a position to look more deeply and act.
> may well catch the eyes and minds of individuals in a position to look more deeply and act.
Of course sometimes that act is "post as many comments as you can to this thread without upvoting it," or "grab about 20 dormant accounts to upvote it, post top-level positive, but dumb, comments, have those same accounts upvote those, and get detected as a voting ring."
Certain opinions, topics, responses get flagged greater now; ones that disagree with a certain political standpoint or that criticize the intel community apparatus get flagged and censored now here. I used to see flagged content where the OP was removed only for bigotry and clearly abusive cases, now its for cases where people feel uncomfortable and strongly disagree.
edit: recently for the first time, I started to feel this is no longer a safe/trusted place for the original hacker ethos it founded itself on.
"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
What big AI companies have done is illegally hoovering up copyrighted creative output of individuals and creating a situation where the wages that normally would be paid to those individuals instead go to that one company (that stole their work) which now becomes disproportionally rich and powerful.
In both cases companies obtain money and power by hoarding information obtained through dubious means (in the former case most academics willingly participate while at the same tone they often don't really have a choice). Exactly what Swartz was fighting against.
I strongly believe Aaron would oppose the appropriation of content. The problem with AI (in this context) is not that the AI companies gain access to information that regular people can't freely access. The problem is that AI erases the information about who originally created a piece of work.
When people want to freely share their work, then they usually reach for the Creative Commons licenses and not for Public Domain, because the latter doesn't protect authorship.
A: "Information is free"
B: "Information is not free"
C: "Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility"
You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you/Swartz to prefer C over B while A is not true?
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws are treated by these companies as mere instructions.
> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend
Big tech will enable access to free information and will help people reach new heights. Do you see how wrong that sounds?
Fun fact: Kim Dotcom is still fighting extradition while these drama queens (I.e Dario) are lecturing us about how much access the peasants should get to AI models fed and trained with stolen IP.
I think what Swartz did was moral, and his prosecution was unjust.
I think what OpenAI did was moral, and them getting sued for it is unjust.
Why do you have one position for Swartz and a different one for OpenAI?
(Aaron Swartz was a mailing-list friend of mine, so I do have some bias here. But in part we knew each other because our moral position on this was similar)
I think many people on HN dont mind OpenAI use of copyrighted material, but do not like how they try to do regulatory capture of a market and try to say their own copyright is now somehow more important.
Like how OpenAI trying to make "distillation" illegal while it exactly what they did with whole intetnet, books, everything.
"Why do you have one position for an activist and another for a eight-hundred and fifty-two billion dollar, for-profit corporation?"
"Why to you have one position for someone who wanted to grow the intellectual commons and another for a corporation trying to enclose it?"
"Why do you have one position for someone who gave his work away for free and another for a company that charges for access to proprietary tech?"
Then, who is "we" here giving the appearance of a consensus opinion here and in mainstream? It's a vocal minority, it's the powerful, it's the causes they fund and put resources behind to continue to preserve their causes. And now today, it's astroturfing, fake AI-LLM-bots almost indistinguishable from you and I. Don't mistake artificial consensus for reality.
> Why is big tech getting away with so much more?
IMO Because the majority of people are passive, standing by, tolerating abuse and trickery by the minority. This is a perpetual cycle in humanity: those minority use their power and leverage and abuse their positions until they are ousted. We have tolerated this because we haven't stood up yet and acted to change things and demand equal enforcement of the laws that appear to apply to us but not to them. If history says anything, they are afraid and panicking and will continue to be more abusive until they push our buttons more and more, and usually it explodes in their face because they still need us (which is why mainstream rich people push robotics and automation and AI down our throats so aggressively because they know all this) and yet never have minority humans won that approach before. Leadership always changes. Life always changes, and no force can stay dominant for ever.
There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
uhhh... capitalism
I do see a problem with a company loudly announcing that they are going to make people's lives miserable purely for profit. Leaving aside that it goes against OpenAI's stated mission ("to ensure that artificial general intelligence benefits all of humanity"), the disdain for the lives they are intentionally trying to ruin makes it a problem.
And even if you believe that the transition is inevitable, as it is the case with phasing out combustion engines in cars, anyone reasonable would see that the transition is gradual to give people time to adapt. Instead of doing that, OpenAI is burning cash at astonishing rates, polluting the environment, and killing personal computing with the only aim of being the only ones left atop the ruins. I do see a problem with that.
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
Judging by how AI threads look like for the past year, they were absolutely right to be worried.
> largest copyright theft operation in human history
In fact, you're doing exactly that right here.
Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
>The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors' copyrights.
Have you looked at their name?
Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
That's also the key political compromise underlying the notion of copyright: that someone is entitled to the fruits of their labor, and should not be economically hindered by a product that could not have existed without said work. That's the basis on which the "derivative work" copyright doctrine emerged: a work sufficiently original that it does not displace the work on which it is based. LLMs fail to abide by that political compromise by a country mile.
I think there's a pretty good argument to be made that every single robot is built on the labor, creative and technical knowledge and advancements of the workers that robot replaced. Robots after all, much more than LLMs, are incapable of creative output. Some human (probably a laborer) figured out how to stamp the steel in just the right ways, or how to cut the patterns in just the right ways so that the product could be manufactured. Then a robot company came in and stole that creative output, or more likely was sold that creative output by the company owners who stole/bought (depending on your point of view about labor and the ownership of labor's creative outputs in the current US legal system) to produce a robot that then displaced the laborer who created the process in the first place.
Plus you somehow didn't even read the article properly, given that "sketchy russian website" is part of a direct quote.
I think they just used a Russian torrent site.
>Microsoft knew about OpenAI’s use of LibGen as early as April 2019
(note: https://z-library.sk/ is prettier/nicer)
you should not enclose the commons.
that is what openai and anthropic have done; capture the commons, lock the model, restrict the outputs.
Further, I'm not sure how they've "captured the commons". By definition the "commons" belongs to us all. Nothing prevents someone else from doing the same thing. That is, unless the Authors Guild succeeds in splitting the courts over the fair use of AI training and the resolution of that split finds that training isn't fair use. Then the commons can only be used by companies or people with pockets deep enough to license the material in perpetuity.
The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
This is some incredible mental gymnastics here, wow. Some books should be public domain (even if they actually aren't), and this magically justifies stealing from an 80TB library of a large fraction of every book ever published including millions of books that have no public funding.
There's a lot of waste associated with preventing distillation. It's a distraction that goes away if we just compel the makers of these trained-on-everything models to publish their weights publicly.
To preserve competition maybe we compromise and give them a three month grace period.
Modify, copy, lease, sell or distribute any of our Services."
IOW, OpenAI declares copying OpenAI Services as either "illlegal, harmful or abusive"
"Services" means ChatGPT, DALL.E, other OpenAI services for individuals, along with any associated software applications and websites
https://openai.com/policies/row-terms-of-use/
https://web.archive.org/web/20260926090525if_/https://openai...
If you look for actual evidence of AI replacing authors the evidence is thin. One study found no real impact, and another found authors using AI (not AI itself!) competing with other authors and applying downward pressure on sales through increased competition. I have seen no evidence yet of AI replacing authors directly, and this is probably their biggest challenge.
But these sound bites sure sound damning though!
I think courts will consider them, but I suspect it will weigh legal doctrines like Fair Use and actual economic studies more than these statements to determine infringement. These statements will likely matter more after a finding of infringement to show willfulness, and hence the damages calculation, if any. IANAL, so a real lawyer should keep me honest here.
I know many people will think that's business as usual, and that business are for profits and that's their only _raison d'être_, that winning the AI race is all what matters, and trusting any promise from a company/CEO/executive make you a fool. But not everyone think like that, and OpenAI/Anthropric/etc. are allowed to exist because many people expect some positive outcomes from their work. And a healthy society needs a bit more than "business as usual/any lie is ok as long as we are not caught" to work. And asking for the same level of responsibility/accountability as any other business would be, in my opinion, a good signal.
And the very first step could be, indeed, to ask them to be accountable here, and the question could be "hacking is about celebrating creativity with computers, creativity is what make life enjoyable, doing creative work is something we can enjoy, why do you want to kill it so much? Why are you so dismissive about people for which creativity is the core of their work? Copyright does not seem to be the target here, as the rent collector is the editor, not George R.R. Martin or any other author, so what's wrong with authors, painters, designers according to you? Similarly, you're only talking about making creative work disappear - not helping creative work in the way it was framed by Steve Jobs - "the computer is a bicycle for your mind". Could you frame your ideal society? What role does culture play in this society?"
I would really like to hear them on these topics. Sincerely.
We are very clever to make it look like the conclusion below follows from the arguments above; but we are also careful to never smudge our bottom line, so it will never change.
the economics of those companies will cause a catastrophic wipe out of jobs across the board.
[1]: https://walt-disney-animation-studios.fandom.com/wiki/The_Li...
[2]: https://jhmoviecollection.fandom.com/wiki/Encanto_(film)/Cre...
Now it is on Hacker News. So what? Are the Chinese models any more wholesome? Are people going to cancel their Anthropic or OpenAI subscriptions and contracts?
One is morally bankrupt. The other is morally consistent.
This is literally a conversation where they’re deciding if they’re willing to take on the risk, then determining “yes.“
And the worst part? They were absolutely right. They have not suffered any real consequences. And when people try to bring up this flagrant plagiarism and theft, they are shouted down by AI evangelists.
Let he who has not downloaded from Annas Archive cast the first stone
[1]: https://www.independent.co.uk/news/world/americas/woman-fine...
The law is not math; intent and outcomes matter, not just the abstract action taken in a vacuum.
It would be an entirely different story if OpenAI was indeed open and the resulting model was available for all. The issue comes when taking information you haven't paid for, and then locking it into a machine you charge for.
Plagiarism is the act of avoiding attribution or citation for personal gain. You can be a staunch anti-IP advocate and still believe plagiarism is ethically wrong.
LLMs are objectively bad at attribution and citation, therefore incur in plagiarism. Even if this is due to a technical limitation, it is still plagiarism.
I want information freed from corporate control, not freed from my control to be gobbled up by corporate interests.
The point was always democratization, not enclosure.
For starters, and this can’t be overstated: scale.
Also, not profiting off it by converting it into a product that competes with the stolen material and its creator. And pirates aren’t generally funded by VC’s.
He gets threatened with a scary letter, his family decides to pay about 3000$ to settle.
Disney has the nerve to run a don’t download music PSA in the form of a Proud Family episode. If you don’t know the Proud Family was a cartoon which attempted to address “black issues”.
Dang hommie, did you know that failing to respect the intellectual property rights of billion dollar corporations is literally worse than selling crack cocaine.
Think of the shareholders! Think of the missed profit projections!
But when billionaires need to effectively resell the IP of all of humanity, that’s just fine.
Since we live in wacky world, Suno which was trained off stolen IP counts Warner Music as one its partners.
The same Warner that was suing over music downloads a few decades ago.
... then again, copyright maximalists REALLY LIKE AI for some reason, even though it's ripping off their property.
"Ripping off" is also doing some heavy lifting, because the judge in the Anthropic lawsuit bent over backwards to keep AI training legal - or so it seems. All the money Anthropic is paying out is for running an internal shadow library, not training books on that library. What makes this judgment palatable to the lawyers is that while AI is stealing a lot of art, it's not imperiling the copyright monopoly. Copyright is a tool that gives artists a monopoly on copies of their individual work, it does not protect artists as a class from competition from non-artists using machines. In other words, the courts are saying, "We know a lot of theft is going on, but we need you to draw the line from a specific individual work to a specific copy".
Patents were invented in Venice to break the power of medieval guilds by making workers trade their collective control over the economy for individual rights to specific inventions only. Copyright was invented by the British crown to reimpose censorship control over printing presses, but it's adoption into American law was based around individual property rights, and thus it has the same problems that, say, Italian patent law has. Namely that it is an artifice to turn a workers right to their labor into a piece of capital that can be traded around like a stock.
This imperils the legal argument against AI training, because the only theft property law recognizes is individual infringements upon individualized property rights. The legal argument against AI training is very collectivist: AI takes a microscopic chunk of every book in existence to create a machine that replaces artists. But copyright only protects the art from copying. Artists are legally unprotected from being copied, and furthermore, the framework of individual property rights that copyright runs on would cause immediate problems if we let anyone individually own the practice of art.
Furthermore, as someone who is part of this "hacker" community, I would like to point out that AI is arguably more corrosive to our norms than to artists' norms. What AI is being used for is primarily satisficing - the practice of giving "good enough" answers, without any of the personal understanding that this community runs on. The ongoing wave of AI-slop decompilations are particularly bad. The assumption with a retro game decompilation is that you take the game apart and learn how it works. The journey is as important as the destination, but when you use AI for this you skip the journey and make the destination pointless.
Of course, management loves this, because their goal is purely to sell you the destination.
There is a parallel set of concerns being voiced by artists, too: that the artistic process is as valuable if not moreso than the actual work product. The underlying principle in both fields is that the labor end wants to learn and develop their craft, while the capital end doesn't care because craft isn't something they can excludably own and trade. We can even see this in the Piracy Wars of yester-decade, or how book publishers and artists react to libraries. Publishers were way more opposed to piracy than artists were, and even moreso for libraries where there's a lot of artists that swear by them as a sales mechanism. The reason why this is the case is that artists can at least theoretically leverage exposure gained from distribution that does not pay them to create more of a market for their work tomorrow. But publishers can't do that - they only buy works from artists and sell them to the public, so once something is in a library or a BitTorrent tracker, it's largely "done" to them.
that was never ok to be done for profit!
Despite it's lofty rhetoric, OpenAI is neither free as in beer nor free as in speech. It is a gang of profit-motivated thieves. Efforts to pirate others works in order to personally profit is the antithesis of the "hacker community" approach to IP. (OpenAI appears to get quite upset when their own work is treated as they have treated everyone else's.[1])
1. https://www.fdd.org/analysis/2026/02/13/openai-alleges-china...
Plagiarism requires near-perfect copying. If you summarize, paraphrase, or re-write something using different language, you're not plagiarizing anything.
Everyone has a right to summarize or paraphrase whatever TF they want. That includes Big AI companies and individuals using LLMs.
It is not worth giving up our rights just to placate a handful of very wealthy authors.
If a dam is built with slave labor, that’s awful and people should be held accountable. It does not, however, affect whether it’s a dam or not.
So call OpenAI "plagiarizers" but not the "devices".
I could imagine taking an existing LLM and telling it “I want to build my own LLM and need as much training data as possible. Go find whatever you can on the internet and store it on SMS://192.168.0.1”. If it downloads TBs of books, did I pirate it?
Doesn't the future look like we are continuing this cycle where the law is only for some people? I'd like to hope not.
It can't be plagiarism, since LLMs do not have ownership over their output. There's no attribution, because an LLM isn't a person.
And to be honest, I do not quite understand the need for attribution, either. Nevermind LLMs, I also don't know or care who made any of the memes I know and share. Neither do I know who wrote vim, grep, Firefox or who invented the jpg format. I don't want to know, either. It doesn't matter.
Maybe someday when you produce something useful that is of somewhat widespread use, you will understand.
But its also a bit funny to read the 2023 people seriously consider that their slop machines will replace writers. Software devs, sure, but writers? Nop.e
They knew the damage they would cause - and went on regardless. The USA need to fix their court system. Right now it is just OligarchBros running the show. The small thief stealing a wallet gets put in jail for years. If you are rich enough, you pay some peanuts and that's the end of it.
Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
It's easier to take over a resigned population. And stating "resistance is futile" is cheap and it might weaken opposition.
Altman has some reason to worry, because developers and the users of HN are then ones he needs to promote/hype his products. OpenAI still needs customers and they need to sell tokens and a large number of users on HN are not only not buying, but are becoming increasingly hostile towards the product Altman is selling. My guess would be that he'd assume that the users on HN are in some sense his peers, and they are currently extremely divided in the question about the value and dangers of LLMs, and are increasingly critical of the business practises of OpenAI (and other AI companies).
It's hard to imagine anything worse than bombing an elementary school and yet people are still using Anthropic products as if nothing has happened.
Optics don't seem to matter much?
> AI agent accidentally publishes OpenAI’s unreleased model weights
The optics are bad. The submission marks a turning page in human history.
We all know through the history that if the benefit of piracy is more than the money, we can't stop piracy.
The purpose of copyright law is to help developing the culture. It means these publishers are doing bad job distributing the copyright works.
The information-centric work is bound to be replicated, plagiarised, generated and anonymized endlessly. It is unjust, unethical and unnatural. But that's what the road we collectively chose in the information age.