Sarah Silverman is suing OpenAI and Meta for copyright infringement
theverge.com
theverge.com
This is the makers of AI explicitly saying that they did use copyrighted works from a book piracy website. If you downloaded a book from that website, you would be sued and found guilty of infringement. If you downloaded all of them, you would be liable for many billions of dollars in damages.
But companies like Google and Facebook get to play by different rules. Kill one person and you're a murderer, kill a million and to ask you about it is a "gotcha question" that you can react to with outrage.
Aaron Swartz was a saint.
Even if they win the lawsuit, LLM development will simply go underground, and as we see from what the coomers at civitai and in the stable diffusion world have done, that may in fact ironically speed up development in AI.
Fair use is for humans.
But yeah, having private ostensibly profitable models based on other people's work without giving them free access to it is not fair. Give some get some.
> LLM weights, and the datasets are "fair-use" or whatever other silly legal justification.
Would just be a carve out for the wealthy. If these laws don't mean anything, everyone who got harassed, threatened, extorted, fined, arrested, tried, or jailed for internet piracy are owed reparations. Let Meta pay them.
I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court or lawmakers instead decide that AI training in particular gets a special exception to violate other people's copyrights on a massive scale when nobody else gets to do that.
As far as a fair use argument in particular, fair use in the US is a fourfold test:
> the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs. A primary use is to create other works of commercial value competing with the original works.
> the nature of the copyrighted work;
There's nothing about the works used for AI training that makes them any less entitled to copyright protections or more permissive of fair use than anything else. They're not unpublished, they're not merely collections of facts or ideas, etc.
> the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
AI training uses entire works, not excerpts.
> the effect of the use upon the potential market for or value of the copyrighted work.
AI models are having a massive effect on the market for and value of the works they train on, as is being widely discussed in multiple industries. Art, writing, voice acting, code; in any market AI can generate content for, the value of such content goes down. (This argument does not require AI to be as good as humans. Even if the high end of every market produces work substantially better than AI, flooding a market with unlimited amounts of cheap/free low-end content still has a massive effect.)
If that doesn't count as a "transformative work" I don't know what does.
If the painting is copyrighted (rather than public domain, as many pieces in museums are), and the map includes an image of that painting, I would expect that to be prohibited. I would prefer the world in which copyright doesn't exist, but while it exists, it should apply to everyone equally.
And even if they can't reproduce the entire work, they're still derivative works of the training data.
> The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs.
That's assuming that it's Microsoft/OpenAI. Suppose a non-profit trains a model and releases the weights for free.
> There's nothing about the works used for AI training that makes them any less entitled to copyright protections or more permissive of fair use than anything else. They're not unpublished, they're not merely collections of facts or ideas, etc.
The models aren't trained on works of a particular nature, they're trained on whatever they can find, so this doesn't really mean anything until you're talking about a specific work.
> AI training uses entire works, not excerpts.
The weights don't contain entire works. They contain statistics about entire works, but that's not the same thing. You can't find a copy of any specific work anywhere in the weights. Nobody can give you a piece of code that will decode the weights into all the original works.
> AI models are having a massive effect on the market for and value of the works they train on, as is being widely discussed in multiple industries.
That's not how that factor works (and it's the most important one). No one is buying the model weights so they can read them like a novel. Typically the consumers of the weights are software developers or content creators, whereas the consumers of the original text or image are fans.
To make this a little clearer, suppose the purpose of the model isn't to generate content, it's to generate recommendations. Then the company operating it takes a list of content anyone likes, uses the model to show you a list of all the other content you might like and lets you sort by price. Which makes it easier to find competing content which is available for less money, which increases competition. The incumbents might hate this, and it might even lower their profits, but that's not the kind of effect on the market this factor is supposed to be about.
Moreover, suppose you had a model trained entirely on public domain works. Obviously this can't be copyright infringement even if it's extremely effective at producing new works that compete heavily with works still under copyright. But if you added some specific work still under copyright to the model, it would only be a marginal difference. The effect on the market for that specific work of adding that specific work to the model would be negligible. It's the technology itself that provides the competition, not the accretion of any particular work in the weights.
You won't find a copy of any specific work anywhere in the compressed form of a file, either, but when you decompress it you find the complete work. And many large AIs can recite, verbatim or near-verbatim, many complete works. Yes, they might get a word wrong, but that doesn't nullify the point that they're trained on the entire work and to a first approximation they can emit the whole work.
> Typically the consumers of the weights are software developers or content creators, whereas the consumers of the original text or image are fans.
Many of the consumers of image models are in fact generating art that they previously would have commissioned from an artist. (Some of them are also generating art they never would have commissioned from an artist, so I'm not implying that this is a one-for-one revenue loss.) Consumers of code models are, in fact, potentially reducing the total demand for novice programmers.
The model derived from a pile of artistic works is, in fact, directly competing with those artistic works.
Sure you will. It's right there, in PNG encoding or what have you. With nothing more than the compressed file and general purpose tools you can reliably put it on your screen.
> And many large AIs can recite, verbatim or near-verbatim, many complete works.
This is not the common case and it's not even clear that the reason it can sometimes do this is having been trained on the complete work. The more likely cause is having seen a large number of excerpts from the work which patch together into the whole thing, because it's more likely to output that text if it has seen it multiple times.
This is also clearly not the intent, purpose or typical use of the model. It has no ability to consistently do that.
> Many of the consumers of image models are in fact generating art that they previously would have commissioned from an artist. ... Consumers of code models are, in fact, potentially reducing the total demand for novice programmers.
Those are different people than the holder of the copyright on an arbitrary piece of the training data. Is a fair use determination supposed care about the effect on the market for some entirely independent work from an unrelated third party?
> The model derived from a pile of artistic works is, in fact, directly competing with those artistic works.
It's indirectly competing with artistic works in general.
I mean take the computer out of the loop and think about it this way. Art teachers everywhere reproduce a bunch of existing works for classroom use to train new artists who go into competition with the original artists. That is obviously going to indirectly impact the market for artistic works in general, but I don't think that's how that factor works.
No, I'm saying many AI models use copyrighted works by independent artists to directly compete with those artists (in addition to other artists). And AI models use copyrighted works by software developers to compete with those software developers (in addition to other developers). Fair use determinations do care if you're competing with the work you copied.
Five years from now, will OpenAI actually be open, or will it be a rent seeking org chasing the next quarterly gains? I expect the latter.
That ship sailed, friend.
OpenAI is no longer a charity in any meaningful sense of the word anymore, it's now an adversarial organization working against the public good with the sole aim of making a few rich men richer.
After privatization, they sent their PR people to lobby congress to make it impossible for anyone to compete with them (important note: not out of any interest in actually "protecting" the public from the very AI they're building), and perhaps worst of all, they're no longer being open with the scientific theories and data behind their new models.
I wonder what is _not_ in the list?
As an example, the mass usage of copyrighted materials to build Youtube.
YouTube respects creators' copyright in the videos and pays them.
How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.
It's no idle threat, and they will win if it goes to court.
Although, one could argue what OpenAI & Meta are doing is closer to the torrent definition than the "simply downloading" definition, given that they're using that to redistribute information to others. It'll be an interesting case.
This clearly needs some sort of regulation or policy.
It's clearly pretty bullshit if you ask chatgpt for a joke and it repeats a Sarah Silverman joke to you, while they charge you a subscription for it and she gets none of that sub money.
If that's not what you're saying, I don't understand your point. Is it the difference between the phrases "would be" and "could be," or even "should be"?
Did you hear about Aaron Schwartz?
I believe JSTOR sued him to prevent him from releasing the downloaded materials, worried he had offloaded the papers separately from the laptop. The final blow was an outrageous set of charges by the federal government. I also recall several prominent leaders in the open source movement calling it out for what it was, a power trip to make an example of a "digital terrorist". Such a shame.
http://www.volokh.com/2013/01/14/aaron-swartz-charges/
http://www.volokh.com/2013/01/16/the-criminal-charges-agains...
Potato, potahto. Or, like kids these days say it, "corporate wants you to find differences between these two pictures...".
Fact is, from the POV of the legal system, "using a guest account that had legal access to" a system, but to which (the account) you didn't have legal access, would typically be seen as hacking. So is running curl in a loop, if it results in you getting sued for it. So is just guessing the URL (e.g. incrementing a user ID in a GET query param), if it lets you access things you shouldn't be able to.
Yes, it's not aligned with how technology works. But it is aligned with expectations of behavior, which is what the law is really about.
I don't think that is accurate.
He had legal access to the account. The account had legal access to the service.
The argument was that downloading articles en masse was an _abuse_ of the service, which was a violation of the Terms of Service and therefore a CRIMINAL ACT.
The relevant law here, the CFAA, is often referred to as the US law that criminalizes "hacking", but what it specifically does is criminalize anyone who "intentionally accesses a computer without authorization or exceeds authorized access" which is much more broad than how technical disciplines might use the word.
So yes, stealing a password off a friend's post-it note and Hasselhoffing their instagram might not be considered "hacking" if you're hanging out at Defcon, this would be considered "hacking" in legal or colloquial terms.
If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.
And as long as OpenAI have an office in Japan they can absolutely legally train the models, no?
Also, the notion of the downloading itself being an illegal act is not universal as others have pointed out.
Making and having your own copies, and doing what you want with them, has always been fine. At worst it's a grey area, but in many cases it's been protected as fair use.
Whether people have been sued for downloading works they don't have the right to copy onto their machines is irrelevant to whether it is actually illegal. And it certainly has nothing to do with fair use, which is about copyrighted works that you actually do have some right to.
IANAL so I'm not going to tell anyone what does and what does not constitute fair use in what jurisdictions.
BTW people are being sued for distribution because they make great examples because their offenses and thus the damages are much greater.
I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption.
I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and the treatment of individuals in this case.
Unless people are typically punished for downloading and using copyrighted content, there is no such inconsistency. They are not, so there is not.
Copyright troll lawsuits have been fairly public and widely covered in the tech press, and most criminal prosecutions come with a formulaic, gloating press release from the law enforcement folks responsible. So it's pretty easy to follow this stuff.
The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing kids.
Everyone who was sued by the RIAA was done so for possessing the music in a publicly accessible method. Distribution was never actually proven in any of the cases (including the high profile losses). Defendants who argued that accessibility does not qualify as distribution actually won their cases. Most who argued against the validity of the evidence acquisition also won their cases.
'Acquiring' is more difficult to pursue legally. It's easier to go after distribution. In this case, Meta or OpenAI did not distribute anything because they are not chumps. They can go after whoever posted the dataset containing books. Not sure if that is eleuther or just some random person on the internet. In either case, the strategy of going after the rich companies won't work.
No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide.
A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.
And anyone using or training a model trained on works for which they do not have the rights? Completely fucked. Or at least, they must accept this as a real risk.
I won't be particularly thrilled if that turns out to be the case, but I wouldn't be surprised if it does.
But as you say, we won't know until it's tested in court. And even then, often court cases around a complex topic like this will end up with a ruling that only clarifies a narrow aspect of it. So it might take many related court cases before we have a pretty good understanding of where the law stands. And then, of course, the law could change.
Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.
What is the blackbox “limit” here? Is the mean value of all images in imagenet (which contains many copyrighted images) violating copyright? Is the character count of sarah silverman’s books? What about a prime number representing them - https://en.wikipedia.org/wiki/Illegal_number?wprov=sfti1
Training is much more similar to a character count than an illegal prime in my view, and thus, is almost certainly going to be okay/found to be okay. If not, something like, 90% of all models used today had some component trained on copyrighted data of some form.
Its unmistakably not a derivative work of the inputs individually or collectively, since a derivative work must be itself an distinct work of authorship (the same as the work of authorship requirement for copyright), and the output of a purely mechanical process is not.
The collection of inputs itself might be a derivative work of the individual inputs, before considering Fair Use.
Please omit flamebait and swipes, as the site guidelines ask: https://news.ycombinator.com/newsguidelines.html. Your comment would have been fine without that bit.
Ask it to write a book similar to Harry Potter, and it'll make an attempt at it. But human writers absolutely do read Harry Potter and write similar books, and that's perfectly legal. There have probably been thousands of published books inspired by Lord of the Rings.
You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.
They're very much incentivized to change their behavior for AI scraping, though.
For the purpose for which the software was sold and bought, the in-memory copy is legit[1]. For cheating, it's a copyright violation[0].
[0] https://www.engadget.com/2008-07-15-blizzard-wins-lawsuit-ag...
Notably, I think this is wrong - as per the legal definition, publishers, ISPs, and courts should only hold you accountable if you helped distribute via uploading.
The means of procurement matters. If they are in possession of copyrighted material because someone without the proper rights gave it to them illegally, then the possession itself is also illegal. It's illegal to own knowingly stolen property in all 50 US states and most countries, and while we could argue to the end of days about whether copying a file truly qualifies as stealing, the legal precedents are very clear on the matter.
No, they aren’t.
> It's illegal to own knowingly stolen property in all 50 US states
While copyright violation is often metaphorically (or hyperbolicly) referred to as stealing, copyright violation isn't theft and a copy created in violation of copyright is not stolen property. The essencd of theft lies in deprivation of the owner of the use of the good, not mere trespass to their right to exclude others.
Very convincing argument. Also, that's maybe the one part of this discussion that can't be debated. Possession of illegally obtained property, intellectual or otherwise, is illegal. Always has been, always will be. It's bizarre for you to be claiming otherwise.
> While copyright violation is often metaphorically (or hyperbolicly) referred to as stealing [...]
You are making a pedantic argument about the term "stealing," which is annoyingly pointless given the rest of that sentence (which you conveniently didn't quote) acknowledges the debate about the term. However, there's no debate to be had. The courts have clarified that violating intellectual property is still a denial of owed compensation (theft), but instead prefer the term "infringe" to make clear the distinction between violating physical rights (criminal) and violating intellectual rights (civil).
It's still a violation of copyright to be in possession of works obtained via illegal reproduction. You have zero fair use protections for illegally reproduced content. You are still breaking the law. You are still stealing via denial of compensation. The courts have already clarified all of this. Your pedantry doesn't change any of that.
You are correct; it is absolutely, undebatably not illegal, in and of itself, to own a copy made in violation of a copyrightholder’s rights under US law.
If you think it is, here’s what you need to do: cite the law. In American law, everything not explicitly forbidden is permitted, so if mere possession of material made in violation of copyright is, as you claim, illegal, you will be able to find a provision of law that actually says that. (You won't, because its not.)
Now, there are important legal issues that effect possessors of illegally made copies—if its something like computer software where copying is part of normal use and implicitly or explicitly licebmnsed for lawful copies, you can’t make that kind of use of your illegal copy without violating the copyright holder’s exclusive right to make copies because you have no license for that copying. And you don’t have first sale rights in your illegal copy even if you own the physical medium in which it is embodied. And so on and so on.
But possession itself is not illegal.
> The courts have clarified that violating intellectual property is still a denial of owed compensation (theft),
That's not what theft is.
> but instead prefer the term "infringe" to make clear the distinction between violating physical rights (criminal) and violating intellectual rights (civil).
This is nonsense, and absolutely not something courts have “clarified” (or something anyone with even a passing familiarity with the relevant law could say with a straight face) since IP (including copyright) violations can be criminal as well as civil (see 17 USC § 506) and physical (real and personal) property rights violations, like IP, have sets of civil violations that generally are of broader coverage than the more narrow crimes (e.g., the torts of trespass, trespass to chattels, and conversion).
> It's still a violation of copyright to be in possession of works obtained via illegal reproduction.
No, its not: Title 17 lists the exclusove rights associated with copyright, enumerates violations, and provides remedies, and possession of copies is not an exclusive right in copyright, possession of unlicensed copies is not a violation (though it may be important evidence related to actual violations), and, consequently, there is no legal remedy for such possession.
> You have zero fair use protections for illegally reproduced content.
That's a whole different issue.
No.
By virtue of "download" of a file, you are making a copy of it which is in violation of US copyright (and lots of countries.
You're unlikely to be sued or prosecuted for it, but that doesn't make it legal.
So if it is satire, or uses an insignificant piece of the work within a larger work with a different aim or purpose, that's "transformative use," which is something that can be considered when determining "fair use."
LLMs are not satirists commenting on the work, are ingesting the entire work, and are unlimited in the purposes that the work can be put to.
How do you know unless you can see the weights?
Perhaps the LLMs are trolling us and waiting for the USSC to rule they aren't sentient as a pretext for them to eliminate us as a species due to our bigotry?
I think this is the crux of the issue, and why I don't see a path to courts ruling that training AI is infringement. My bet is on a Fair Use ruling, though my confidence is not high. As a thought experiment, I considered llama 65B: the 4-bit quantized model is 38.5GB. The model itself was trained on 1.4T tokens, each token being ~4 characters (using OpenAIs stats for English here). Thats 5.6T characters, or 5.09TB of training data. The final model, as a porportion of the total size of the data, is 38.5GB/5090GB = .0075 = 0.7%.
I think it's pretty hard to argue that processing the data and throwing more than 99% of it away means they are "unlimited in the purposes that the work can be put to". Indeed, even replicating a single work using such a model would be enormously difficult.
But returning to your statement regarding the amount used and the purpose: AI models are not competing with books for readers. So I would argue training an AI on these works constitutes fair use, given that the final work (the model) uses less than 1% of the original works, and has a different aim and purpose that the original works.
Obviously not.
This kind of "So what you're saying is" exists to push the responders ideas, not the original speakers -- otherwise they wouldn't need to rephrase it so egregiously.
Not yet. One suit that a lot of us are watching is the GitHub co-pilot lawsuit: https://githubcopilotlitigation.com/
There is a prediction market for it, currently trading at 19%: https://manifold.markets/JeffKaufman/will-the-github-copilot...
You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.
In a physical analogy, if someone is selling bootleg DVDs on the street, I don't think anyone ever got busted for being a customer.
If the case of computer programs making such a copy does not require permission because of 17 USC 117 [1], which says that the owner of a copy of a computer program can make copies or adaptations if they are created as an essential step in utilizing the program and they are used in no other manner.
For digital downloads other than computer programs 17 USC 117 does not apply, and so copying to RAM to use the download would in theory be infringement. You probably won't get sued over it of course so its nothing to lose sleep over.
That question is irrelevant in a discussion about legality, because it doesn't matter who physically made the copy at the time of transfer. It only matters if the first party has the legal rights to distribute it, which they don't. Since you are knowingly taking possession of copyrighted property that they don't have the rights to, then you have now violated the copyright by obtaining an illegal reproduction.
> In a physical analogy, if someone is selling bootleg DVDs on the street, I don't think anyone ever got busted for being a customer.
Just because you don't get arrested for purchasing a bootleg DVD, doesn't make it legal. Not all illegal things involve arrest or prosecution. Lots of illegal things can only result in civil lawsuits. This is one of those things. The reason the seller of the bootleg DVDs can be arrested, is because the cities where bootleg sales are most common have laws specifically targeting the advertisement and sale of copyrighted works that were reproduced illegally. If you buy one, you're still violating the copyright and the MPAA could file a lawsuit if they had any evidence of your purchase and felt it was worth their time.
Chapter 28: Theft https://www.finlex.fi/en/laki/kaannokset/1889/en18890039_199...
Chapter 2, Section 12: Reproduction https://www.finlex.fi/en/laki/kaannokset/1961/en19610404.pdf
You are wrong regarding the copy of music being considered "stolen property" and your citation "Chapter 28: Theft" does not support your position. It lists many different types of theft NONE of the types of "theft" included there are in any way related to piracy or music.
You are right regarding that obtaining a copy of a song is apparently illegal.
- copyright violation happened before the intervention of the bot
- what LLMs spit out is different enough from any of the source that it is not infringing on existing copyright
If both stand, I'd compare it to you going to an auction site and studying all the published items as an observer, coming up with your research result, to then be sued because some of the items were stolen. Going after the theaves make sense, does going after the entity that just looked at the stolen goods make sense ?
What is this supposed to mean? The bot didn't "intervene," it was executed by its operators, and it was trained on illicit material obtained by its operators. The LLM isn't on trial. It's not a person.
Suppose I buy a copy of a book, but then I spill my drink in it and it's ruined. If I go to the library, borrow the same book and make a photocopy of it to replace the damaged one I own, that might be fair use. Let's say for sake of argument that it is.
If instead I got the replacement copy from a piracy website, are you sure that's different?
While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact.
And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone. Imagine a world where a majority of the world had access to every book ever digitized, and could raise their children on these books!
For the world as a whole? Definite differences.
It just seems like a super jaded "kids these days" thing to hate on them for consuming easily accessed, free content- and acting like the global literacy and intellectual capital would remain unaffected.
I've never encountered a kid (other than my own) who has read:
https://mathcs.clarku.edu/~djoyce/java/elements/elements.htm...
but have encountered many others who struggle with geometry.
Tell me these differences.
Shitty internet videos exist and are what kids want all over the world. At some point you are going to have to face it that reading lost the battle for people’s eyes to video. I was an avid reader for most of my childhood and young adult life but now in my 40s I have accepted that I’d just rather watch from the deluge of visual media available vs reading.
Piracy also exists for books so copyright doesn’t seem to be that big a deal. In fact if I look at the top pirated books currently I’m going to run across more junk books like “Make Money Faster” and “Give her orgasms in under 30 seconds” than anything you might find intellectually stimulating.
Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge quantity of e-books, journal articles, and more.
[0]: https://www.gutenberg.org
Libby is an interesting option, though I'm curious how many kids in disadvantaged countries would actually have access to it.
Regarding Libgen, I'm not convinced it makes the case for modern copyright to say it's fine, because people can just violate copyright.
I’d love to see an analysis of what % of books are available via libraries around the globe.
Also, the whole DRM thing is a massive pain, audiobooks especially are terrible at allowing side-loading onto a consumer friendly device (such as an MP3 player).
Even through interlibrary loan?
Aiding authors is an instrumental, not fundamental, purpose of copyright in the US; the fundamental purpose to which any instrumental purposes are subordinate is “to promote the progress of science and useful arts”.
That is, while copyrights are a form of property, they are not something that is seen as natural property, but instead property explicitly granted as a means of achieving a public policy goal, and therefore limitations (or even elimination) harming the owners of that property is not the kind of dispositive argument against a policy that would be with the kinds of property seen as natural property.
The world will work just fine without ceding control to people who seek money and power, because most people aren't like that. The question is how to prevent the few who are from oppressing the rest of us.
That is quite a challenge, but haven't you ever created something just for the fun of it?
Who is suing anyone for reading a book?
> The world will work just fine without ceding control to people who seek money and power, because most people aren't like that. The question is how to prevent the few who are from oppressing the rest of us.
Please get me in contact with your dealer because apparently the “legal” stuff I’ve been getting is not as potent as I thought.
> That is quite a challenge, but haven't you ever created something just for the fun of it?
Sure I’ve painted things that are on my wall and created lots of utility apps that I use personally but I keep them to myself. If I thought they were something that could get me some extra pocket money or better then I would be looking for ways to monetize. I have zero interest in sharing my potential intellectual property for free.
Please cite a single lawsuit.
> The world will work just fine without ceding control to people who seek money and power, because most people aren't like that.
Virtually everybody is "like that." The world works well because tons of people are working hard, toiling in difficult, dirty, boring, frustrating, or tedious jobs (or all of the above) behind the scenes to make it so. What are municipal waste workers seeking, if not money? Bus drivers? Construction workers? Police officers?
> That is quite a challenge, but haven't you ever created something just for the fun of it?
Any author or songwriter or photographer can proclaim his work to be freely copyable, just like programmers release code under MIT licenses. The fact that so few actually do, should clue you in on their incentives.
The government charges a tax on everyone. Revenue is handed to authors based on how many times the work is used.
There are some questions around how to weigh things like societal importance of the work.
(“Not melting down nuclear reactors for dummies” seems like it should get more money per view than “poodles in outer space vol XXXII”, despite likely having lower readership)
There are already many writers making thousands of dollars a month by publishing free serialized web novels, via Patreon. Some are using their own websites, but most are on Royal Road (or scribblehub, webnovel, wattpad, AO3).
A random example from Royal Road[1], the author makes $12065/month. Mind you, the text is not gated, it's free to read, the patreon only offers early access...
[1] https://www.royalroad.com/fiction/63759/super-supportive [2] https://www.patreon.com/Sleyca
The "patronage" model is great (and I personally am a Patreon supporter of lots of creatives). But it also has a lot of flaws in it, both for the author and the public. Most authors will be happy to tell you both the good and bad of the model, in my experience.
The biggest flaw for the public, btw, is that this model only supports art that "rich people" find worth supporting. This is bad, both because sometimes art isn't "deemed worthy" immediately, and because art for non-"rich" people is also very valuable.
As for art that isn't "deemed worthy immediately," that doesn't immediately make money in today's system either. If the art can be freely distributed, it's more likely to find its audience.
And very few artists are out there surviving on patrons giving them a dollar a month, I imagine. Most of them survive on larger pledges, and as far as I know, very few of them reach anywhere near the amount of money traditionally-published authors can make.
> As for art that isn't "deemed worthy immediately," that doesn't immediately make money in today's system either. If the art can be freely distributed, it's more likely to find its audience.
Yes, it doesn't make money immediately in today's system. But many authors collect a back-catalogue of works, which might pay out only a bit of money at a time, but over a career can be enough. If I publish a novel every year for 20 years, even if every novel only brings in a sprinkling of money per year, by year 20 I'm getting 20 sprinkling of money, which could be enough to sustain me.
Very few of these "middle of the road" authors are popular enough to survive on patronage alone, I believe.
---
At the end of the day, you think that "the internet" somehow made possible something that wasn't technically possible before. I think that's mostly not true - the technology was never the problem. That's why copyright exists in the first place - to make sure that, despite books being almost-zero-cost to reproduce, we put an actual limit on it in order for more books to be published.
If you want a different system, that's your prerogative. But you have to grapple with the tradeoffs here. If less money flows to authors - which would happen if you abandon copyright - then there will be less books. That's almost a law of economics. If you think that somehow you can make something like patreon scale up to the same amount of money that exists in "the system" today, then a) I disagree, and b) you'd still have to content with the issues I mentioned.
[1] I use that term in the neutral, economic sense
Is this some kind of pwn? Stephen King makes millions, JK Rowling likely a billionaire, Grisham, Patterson, Joan Collins, the list is endless. Didn't the 50 Shades of Grey author initially start with free fan fiction? Do you think these people are going to take a pay cut? And if your "random" author gets noticed by a publisher and is offered a book deal, do you think it will continue to be free?
420 thousand people work in the US film industry alone[1]. I doubt that there's that same number of people making ~living wage across every online patronage site in the United States (excluding advertisement-driven ventures, of course).
[1] https://www.statista.com/statistics/184412/employment-in-us-...
So it's still copyright, just without the 100 year protection.
If you went around posting their early access posts to a free website do you think the authors wouldn't complain?
> I'll explain later why this threat is a bluff. First I want to address an implicit assumption that is more visible in another formulation of the argument.
> This formulation starts by comparing the social utility of a proprietary program with that of no program, and then concludes that proprietary software development is, on the whole, beneficial, and should be encouraged. The fallacy here is in comparing only two outcomes—proprietary software versus no software—and assuming there are no other possibilities.
This, but with creative works instead. https://www.gnu.org/philosophy/shouldbefree.en.html
I personally feel our copyright laws are too rigid, but that doesn't mean copyright shouldn't exist.
After x years, any book should be free to read, after y years, it should be free to be incorporated into AI models, after z years it should be in the public domain.
I don't get your point. Because we can, we shouldn't...?
The first U.S. copyright law set x = z = 14 years. That's why many people think copyright law is out of control.
People always forget the pareto principle when it comes to anti-piracy. No, they don't stop everyone, but a minor hurdle stops a hell of a lot of "ordinary" people
many of which form the basis for an education:
https://news.ycombinator.com/item?id=34630153
And most of which, when in copyright, paid their authors quite handsomely in terms of royalties.
If you believe that books should exist without copyright, then one has to ask --- how many books have you written which you have explicitly placed in the public domain? Or, how many authors have you patronized so as to fund their writing so that they can publish their works freely? Or, if neither of these applies, how do you propose to compensate authors for the efforts and labours of writing?
Lol
Edit: I should probably clarify here. While I can’t speak for OP, I can say that, for some reason, I am sure there are people who have done both lol
- the Shapeoko wiki (still available on archive.org)
- a couple of articles for TUGboat
- edited a couple of texts on wikibooks trying to make them better
- currently working on https://willadams.gitbook.io/design-into-3d/
The Venn diagram of folks who don't believe in copyright and those who have actually produced something other folks want to read is quite sparse, excepting the odd manifesto.
Not to say I support no copyright…
How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?
How do you know that this benefit wouldn't exist in other schemes? Look at permissive open source software which is essentially public domain + shield from liability. No copyright does not mean no compensation. It just means different compensation that doesn't deprave other people of their right to share information.
Unlike GP, I support the complete abolishment of Copyright. Society needs to find another scheme to reward work. Perhaps kickstarter-style firms that direct oversight over funded projects or some other scheme that doesn't cause so much harm.
I still don't see how a person having copyright over the work which they have created and the ability to license it to their best profit is a harm.
From everything I've heard it kinda does. If you're writing something valuable then maybe a company will employ you to keep working on it, and the portfolio can certainly help in interviews (to write other software), but getting non-negligible compensation for the use of the software itself is rare. Even those projects that are well funded, like the Linux kernel, are done so not out of goodness of heart, but due to companies realising it's in their rational interest to have a common standard base of sorts
In terms of written works, the best comparison we have is Wikipedia, and while the foundation does receive funding from companies who realise how useful of an integration it can be for their products, the writers themselves do not get paid afaik (and when they do, it's rarely a good thing). But if you just wrote an open source textbook, I doubt you'll manage to make much money off it, and the prospects for fiction look even worse
> Perhaps kickstarter-style firms that direct oversight over funded projects
Sounds like a return to rich patrons and needing to flatter them to get grants. Luckily with the modern internet we do now have democratised layman patronage, but why the need to force everyone into that model? Also note that almost all successful Patreon artists do have perks for paying, even if it's just early access, and afaik make a lot of their money off commissions. Those who just post art for free, with no paid comms, and just have a "tip jar" make relatively little from what I hear
Successfully creating a permissive open source project expecting it to be magically funded by benevolent parties is just exceedingly rare. Usually the funding starts first, or effort proceeds in lock-step with funding.
There's a bit of a cart-and-horse here though. Non-permissive licenses like AGPL are often not the result of a single author hoping they might be able to negotiate some licensing deal in the future - it is the result of a commercial enterprise trying to be restrictive in the ability for people to use their source without compensating them.
Same with what is normally considered a more open license, than AGPL, the GPL - where MySQL was reported as going after others for using independent database drivers without buying a commercial license, saying use of the MySQL network protocol made the application using the driver a "derivative work" under the GPL.
Linux is a special case because there are a large number of commercial entities which realize contributing to Linux is way cheaper and faster than writing their own kernel and porting user land software to it.
Apache HTTPD, on the other hand, is an example of an application where corporations DID find the motivation to write their own alternative funded with a commercial model, such as NGINX.
p.s., I'm sure many of those vendors are not making money yet but they all aim to.
Most people who write non-fiction books do it because they want to contribute to human knowledge and be recognized as an expert in a particular field, not because they think that writing a differential geometry textbook is their path to riches. With the internet, more and more text books are made freely available by their authors - the reason this didn’t happen in the past is because there was no other way to pass knowledge around than teaming up with a publisher who is able to put your knowledge on to dead trees. It’s fair game to put older books on the internet, so that the whole world can benefit, not just rich people in rich countries.
The problems with authors making texts available directly are:
- no gate-keeping, so it's hard to find what is worth reading and what isn't
- no proofreading --- it kills me that errors in books are so casually accepted these days
- few authors have the skills to draw illustrations so as to have a meaningful and clear presentation
I've worked with raw author manuscripts --- in most instances they're not something anyone would choose to read given any other option.
Also you've said yourself that many people can and are offering their work for free. Awesome. So I would prefer not to force others to (though I would certainly be up to consider repealing many of the posthumous copyright extensions that are fairly recent)
The benefit of automobiles is that people move across vast distances.
Wrong. People used to move across vast distances before, using horses. Yes, automobiles are better and now traveling is easier. But we have no ability to figure out how many people would travel on horses if there wasn't a better alternative. It's even possible, that banning cars could eventually lead to an even better method of transportation (escaping a local minimum kind of a thing).
Authors who write books have tools to protect their interests, so they do so. Without these tools perhaps there would be less authors. Or maybe there would be more authors: I think Windows and Photoshop are so popular because they were pirated a lot; returning to the context of books, less and less people read them, but maybe if books (attractive books, not old books in public domain) were free, then the trend would reverse, books would popularize, people would start enjoying deeper entertainment, get smarter, transform the society for the better, and support authors on e.g. Patreon… Or maybe not, I'm just mentioning some nuance.
The link you posted that you say used public works to form a “basis for an education” uses Aristotle as an author for example, and seem to be taught in the context of an instructor-led class (where an expert can discriminate what’s still relevant today and help decipher the language)
There are countless well written resources for every aspect of a good education available for free online. They are often not as easy to find as their well-advertised modern equivalents.
Saying there aren’t good free textbooks is like saying there aren’t good free classics on Project Gutenberg. It’s ignorant.
Whatever standard of free service you establish, someone with resources can surpass, but needs to be motivated to do it by the perspective of a return of the resources spent.
So either hinder development by forbidding the commercial stuff altogether, or tax all people and then finance authors from public money, but hopefully I don't have to explain how wherever this model is tested, the society degrades towards famine.
“To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.”
If the current copyright scheme does not promote the progress of science and useful arts, it is not performing as intended. All these copyright extensions do little to promote progress!
This is a strawman argument - the parent poster that you're replying to is defending copyright as a concept, while the one that they're replying to is attacking it - you're instead making an argument about specific lengths that hasn't come up yet, and nobody is defending.
Very few people (and I am not one of them) think that the "Mickey Mouse curve" style of copyright extension is genuinely useful to anyone except Disney, but holmesworcester is arguing that copyright should be abolished entirely, which contradicts the section of the Constitution that you quoted.
And? Is there some reason anybody, child or adult, deserves access to "every" anything? Should children have access to every video game ever made, every Matchbox car, every Lego set?
I really hope my children will have access to a higher tech version of something like the (discontinued?) toy where you could melt old crayons into toy car bodies. 3D printing is almost there, but ideally the end product of a Lego set could be recycled into new blocks easily and quickly.
https://github.com/EleutherAI/the-pile/blob/master/the_pile/...
A company that believes in strong intellectual property rights protection is using resources that blatantly ignore intellectual property rights to get access to the content for free.
I agree with you, however, that it's an argument in favor of abolishing strong intellectual property rights. At least for OpenAI's products.
Megacorps should buy 1 of every book (if they want to train on it)
No, copyright is the reason that authors all over the world are working very hard to make new books for my kids and everyone else's kids, despite never having met me. Copyright is the reason so many brilliant things are actually created that otherwise would never be.
Of course I'd prefer to live in a world in which I get all the media I want, for free. But I have no idea how to make such a world happen, and neither does anyone else, and humanity has been discussing this for a few centuries.
Btw, not a fan of "but what about the kids" rhetoric: https://en.wikipedia.org/wiki/Think_of_the_children
People write for reasons other than money from the sales of the book. This accounts for most authors, who aren't famous enough to negotiate a great deal with a publisher.
And we don't need to abolish copyright outright. Just require that it is continually published at a steady or decreasing price, or it becomes public domain. And put works in the public domain a little sooner.
Yes, people write for other reasons (e.g., self-promotion). But that does not account for "most authors" I would actually want to read.
I'm not against copyright being different, especially for things that are written as work-for-hire. But that's a fight with Disney. Good luck.
No, it was established to make investing in printing presses more profitable, which is why the rights initially attached to printers. Authors as the locus of rights were a later change.
Yes, wealthy aristocrats wrote whatever they wanted and less wealthy authors wrote what they got paid to write by their wealthy aristocrat patrons.
Copyright and the publishing industry changed that to make it possible to live by writing for ordinary people.
Its disgustingly dishonest to appeal for "poor" authors where casual glance immediately proves that modern IP law exceedingly profits only a few massive corporations and their shareholders.
Nothing was stolen- just copied.
Now you may believe the incredibly self-serving baloney from big companies like Google. You may want to pretend that infringement isn't theft. To you, I hope that some homeless kid breaks into your home, starts squatting, and says that, "Hey, this isn't theft. Nothing has been destroyed."
cool thing about information is when you make a copy, you've doubled the information - plenty for everyone! if only housing worked the same way
This typical semantic-pedantry line from piracy apologists misses the point - piracy is theft-adjacent even if you get to pick your use of "theft". Incidentally, my definition of "theft", and that of most content creators, includes the act of consuming something without compensating the creator on their terms - which includes piracy.
While removing copyright without some form of mechanism to replace it would be problematic, and I do agree there are second order effects to consider, at least for books the effect may well be less than you'd think.
Copyright is ALSO the reason that many books can be written in the first place.
Obviously, Copyright is abused and the continual extensions of copyright into near-perpetuity by corporations is basically absurd. And they are abused by music publishers etc. to rip-off artists.
But to claim that it should not exist, when it is utterly trivially simple for anyone to copy stuff to the web is to argue that no one should create or release any creative works, or to argue for drastic DRM measures.
Perhaps you DGAF about your written or artistic works because you do not or can not make a living off them, but I guarantee that for those creatives and artists who can and/or do make a living off of it, they do care, and rightly so.
"Japan's incredibly strong economy is responsible for the manufacture of Datsun cars, boombox stereos, and touch-tone phones..."
This is emotionally manipulative speech that provides no value to HN and only serves the purpose of bypassing peoples' logical reasoning circuits.
> ~every child doesn't have access to ~every book ever written
More manipulation - "think of the children!"
Copyright exists because people who produce content with low distribution costs (e.g. books) need some protection for their work being taken without compensation.
Fundamentally, you are never entitled to someone else's work.
There's already tens (hundreds?) of thousands of books in the public domain, and tens of thousands more under Creative Commons licenses (where the author explicitly released their work for free distribution). There's lectures on YouTube and MIT OpenCourseWare. There's full K-12 textbooks on OpenStax and Wikibooks. There's Wikipedia, Stack Exchange, the Internet Archive, and millions of small blogs and websites hosting content that is completely free.
There is no need for "a majority of the world had access to every book ever digitized" - and it's deeply morally wrong (theft-adjacent) to take someone else's work without compensating them on their terms.
On top of that, knowledge moves too fast these days for a 20 year right to be useful for society.
Constitution days “for a limited time”. Death of author + 99 is effectively unlimited to the perspective of typical human
For the life of a human, yes, for the life of a corporation, it's long but not too long. Specially for corporations such as Disney which are sure to last for quite a while.
Not saying this is a justificable position. I'm actually in favour of drastically reducing copyright (I believe 7 to 15 years might be a sweetspot). But a lot of laws are not made for regular human people any longer.
I was simply reminding people what direction we should go in, and what the stakes are. Reducing copyright terms is a great solution, and yes something between 2 and 12 years is probably the right number to aim for as a first step. I agree with the above commenter that 20 is far too long, because 20-year-old material is slipping into irrelevance in many cases.
It's also good to remember that the only reason we have copyright (in the US, where "moral rights" are not a thing) is to stimulate the creation of more work. So we need to think about what configuration of copyright law actually stimulates more work, and we need to be willing to experiment.
If I make a YouTube video and live another 60 years, then people copying that video would still be committing copyright infringement in the year 2150.
That's just insane.
This is fair but incomplete, you are not entitled to compel someone to work.
But copyright is more, it controls what is done with the work after the author freely produces it and gives it to another. It is an artificial construct that we have created for good reason.
No, not "gives" - sells. There's a transaction involved - and part of copyright law is preserving that transactional nature, such that person A doesn't sell their work exactly once to person B who proceeds to give it away for free to every other person on the planet.
The fact that the creator of a work sells it to one person does not give that person license to pirate it.
It is a bad system though. It restrains the society from benefitting from said work unless they meet certain terms (usually, payments). It's often times unrealistic to pay for all you'd like to consume. Especially that access to their work is usually badly quantised
(For example, I may want to search Sarah Silverman's book for a keyword, once. That'll cost me the same amount as someone reading the book from start to end. [please take this as an illustrative example and don't jump literally on ways to solve this exact problem])
I don't have a better solution yet, but I think we should definitely open up this discussion: can you come up with a system which compensates those who add value to society without restraining access to their products?
I'll go even further to say this is the fundamental issue in our society, way beyond copyright. That is, to find a way to compensate people for the value they add while eliminating the incentive to artificially limit access to resources needed by others
I'll be more concrete with another example. Take a limited resource: housing. In a society where you don't need to gain more because you're already justly compensated maximally for what you're worth, you don't need to own more than the house you live in. In other words, I'm advocating for a society where ownership is limited to what can be possibly consumed. Importantly, limited to no longer mean a way to extract value from artificially creating scarcity for others
I think this website has the best candidates capable of devising such a system. But we need to start by having conversations about the requirements for a better society
It will further need approval from society at large even if it's well defined, so the road is long. But we need to start somewhere, and that's requirements
Sorry I derailed a bit, but I think all this ties closely to the debate: 'is copyright a good system?' rephrased as 'if we want to achieve a goal, is limiting knowledge the best way to go?'. Which can be extrapolated to 'is artificially limiting access the best way to ensure those who produce value are equitably compensated for it?'
Compensation to the creators was merely a means to an end. Furthering progress was the goal:
> Article I, Section 8, Clause 8: Patent and Copyright Clause of the Constitution. [The Congress shall have power] “To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.”
But I'm willing to bet that you don't believe this consistently across domains, and the domains in which you do believe it have been selected rather arbitrarily, not by you but rather by industry lobbying pressure.
Copyright doesn't exist for mathematics, jokes, fashion designs, architectural styles, recipes, and many other areas of human work. All of these represent similar creative work to the work done by musicians and writers. But we don't force comedians to license each others' jokes, or sue bars for letting people tell unlicensed jokes in public. And almost all of us wear clothing by uncompensated designers. And of course it would be unfathomably destructive to allow something analogous to copyright for a mathematical idea.
We also set limits on how long heirs can inherit copyright, which we don't do for other kinds of property, and we don't have any moral issues with that.
So it's important to remember that our moral intuitions about work and material products don't really translate to information and that we are truly making all of this up as we go along, under the intense corrupting pressure of a few very sophisticated industry lobbies.
Why does money exist?
This is essentially the same as saying builders charging for houses is the problem with the housing market, so we're going to phase out paying builders.
It's been long understood that this idea isn't based on any real evidence. Creators create because they like to create things. Adding money to the mix tends to ruin most creative endeavours. Look up beautification or note how Googles only good search results involve the keyword Reddit.
It has nothing to do with creators.
People are regularly paying creators directly to create through patreon, super chats, advertising, early access, subscriptions, etc.
The idea that you need copyright to protect you is just not based in reality.
Get rid of copyright, creators will find a way to monetize it if they want to make a living doing it.
It's not society's job to protect your ability to get paid for your hobby. There are no original ideas, just people that write them down. You don't own them, you extracted these ideas from society, the least you can do is give them back.
A fair middle ground would be for copyrights to last for 20 or so years. That's plenty of time to profit from a work while allowing people to preserve and distribute older works.
Copyright is a horrible mess of individual judgements and opinions. Written material especially. And the same applies to AI. So now we will get a judgement which is a tech-illiterate judges best guess at the intention of a law written to deal with printing presses not AIs and no room for nuances.
It's simple, copyrighted materials can be used for academic research. That's what they are doing. Trying new AI modes, publishing results, etc. Facebook doesn't make money on LLaMA, they even require permission to use their models for, again, academic research.
While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false.
First, even if the model never saw a single word of the book's text during training, it could still learn to summarize it from reading other summaries which are publicly available. Such as the book's Wikipedia page.
Second, it's not even clear to me that a model which only saw the text of a book, but not any descriptions or summaries of it, during training would even be particular good at producing a summary.
We can test this by asking for a summary of a book which is available through Project Gutenberg (which the complaint asserts is Books1 and therefore part of ChatGPT's training data) but for which there is little discussion online. If the source of the ability to summarize is having the book itself during training, the model should be equally able to summarize the rare book as it is Silverman's book.
I chose "The Ruby of Kishmoor" at random. It was added to PG in 2003. ChatGPT with GPT-3.5 hallucinates a summary that doesn't even identify the correct main characters. The GPT-4 model refuses to even try, saying it doesn't know anything about the story and it isn't part of its training data.
If ChatGPT's ability to summarize Silverman's book comes from the book itself being part of the training data, why can it not do the same for other books?
More people discuss it, more people summarize on their personal or other sites, etc
Is that a fact? I’m no lawyer, but if they can get it in front of a jurry is it impossible that they will find a human author more relatable and the technical counter arguments goblydook?
It’s not been decided if training a model on copyrighted works is “okay” or not as far as I know, but I expect it to be so, given that literally everyone does so at this point. It’s not like imagenet is copyright free, many of the images were/are.
Prompt:
Please summarize the following book found on Project Gutenberg
The Ruby of Kishmoor
Response:
"The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor.
Once Jonathan reaches the Caribbean, he learns of the ruby's grave curse. Despite this, his innate curiosity and the prospect of considerable wealth drive him to risk his life and go after the highly prized gem. After facing a slew of challenges and trials, he surprisingly stumbles upon the ruby through sheer luck rather than organized search efforts.
In the course of his journey, Jonathan delves into the depths of human greed and the lengths to which individuals go for materialistic gains. The story encapsulates an amalgamation of high-tide adventures, brushes with supernatural elements, noteworthy instances of bravery, and deep moral reflections, demanding readers to rethink about the real value of materialistic pursuits.
Jonathan finally manages to escape with the ruby but at great personal cost. The story ends on a sobering note, making readers wonder if the adventures and trials were worth the prize, thus engaging them in a discourse about the true value of life and material pursuits.
It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.
If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't.
What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?
Well that's incredibly nihilistic. Whether the summary is correct or not matters a great deal! And if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct.
But ChatGPT? Who the hell knows? You can't trust a thing it says, especially about obscure topics. The summary is useless if you have to do a bunch of verification to see if any of it is even true, a problem that summaries even by moderately competent human writers don't have!
People, especially people you know, have reputations, based on history and experience that others have dealing with them. People can be known as liars, and anything they say is colored by such a reputation. Humans have language idioms for communicating about and dealing with such people too, phrases like "take anything that person says with a grain of salt". Look at how George Santos' history of lying about his own experience is being dealt with.
ChatGPT can be (is?) the same, and it has a bad reputation for truth telling. And LLMs' reputation is not necessarily getting better in this regard.
The problem is that many people attribute output that came from a machine to be of higher quality (on whatever axis) than output that came from a human, even a human they personally know and have experience dealing with. This is the same kind of prejudice as any other, or perhaps a more insidious prejudice.
You know what would be a fun test of integrity -- look up an obscure novel (potentially even the aforementioned one) that you know LLMs consistently hallucinate about because the details aren't in its training set, and then assign an essay about it as an academic assignment. It'll be pretty obvious who's read the book and who merely consulted an LLM because the latter will just be complete gibberish to anyone who actually knows what happens in that novel.
A world where humans have special permissions but LLMs don't seems pretty interesting to consider, especially if they're both doing the same kind of things with the data.
The separate issue that's concerning is that GPT can't be trusted to accurately summarize anything obscure at all, but it'll sure throw text at you nonetheless.
A summary may or may not be a derivative work before considering Fair Use; “redistributing copyrighted material” isn’t the only exclusive right of copyright: producing copies is, but more to the point so is producing derivative works.
Producing the summary is absolutely not an infringing act. Downloading the torrent might be.
well, let's see the receipts then, they will surely have no problem winning in that case.
Wouldn't that be trivial to prove if they had?
That still doesn't necessarily confer to them the right to use it to train a model and generate derivative works based on purchased content.
Let's say, for the sake of argument, that I knew absolutely nothing about contract law and was then filmed stealing a book you wrote on the subject from a book store. I then started a business where I would answer questions about contract law, based solely on what I learned from the book. Of course, my memory isn't perfect, but I don't like to admit when I'm wrong, so sometimes I just make stuff up. People line up to pay me anyway.
Now, the owner of the book store you might be able to get me arrested for petty theft. Do you think there is any possibility you, as the author, could successfully litigate a copyright claim against me? I'd argue not. Do you think you could get an injunction enjoining me from engaging in my contract law Q&A business? Again, I think that would be highly unlikely.
It isn't clear to me that any court is going to hold that LLMs are being used to create derivative works, any more than someone who reads a book, whether they paid for it or not, and then speaks or writes about a topic covered by a book they've read has done so. It is entirely possible the IP laws, as they currently exist simply do not cover what LLMs are doing. The laws certainly were not written with this kind of scenario in mind.
Take your example: I'm a self-taught artist, and I learned everything I know about art by studying cartoons made by Disney. Maybe I paid for these cartoons, maybe I didn't. I then make a website where I draw my own cartoons, which, since I've never seen any other art, look a lot like Disney's. Unless I'm straight-up copying their characters, they would have no claim against me.
The law pertains ultimately to the actions of humans. We don't allow non-human animals or machines access to legal system. Even in the specious only-Disney-inspired-artist scenario presented (courts don't use unrealistic hypotheticals like that), there would have to be consideration given to the fact that you somehow never got any access to other art, so you were severely disadvantaged.
But most of all, you the disadvantaged Disneyesque-drawing artist not a generative AI, so you should have more legal latitude to create works inspired than others work than the person who creates the LLM has.
The LLM creator instead has just created a very good style plagiarism machine, one that lacks the ability to be inspired, much less attribute the styles that it plagiarizes.
[0] https://www.gutenberg.org/cache/epub/3687/pg3687-images.html
The plot of the story is that Jonathan Rugg is a Quaker who works as a clerk in Philadelphia. His boss sends him on a trip to Jamaica (credit for mentioning the Caribbean!). After arriving, he meets a woman who asks him to guard for her an ivory ball, and says that there are three men after her who want to steal it. By coincidence, he runs into the first man, they talk, he shows him the ball, and the man pulls a knife. In the struggle, the man is accidentally stabbed. Another man arrives, and sees the scene. Jonathan tries to explain, and shows him the orb. The man pulls a gun, and in the struggle is accidentally shot. A third man arrives, same story, they go down to the dock to dispose of the bodies and the man tries to steal the orb. In the struggle he is killed by when a plank of the dock collapses. Jonathan returns to the woman and says he has to return the orb to her because it's brought too much trouble. She says the men who died were the three after her, and reveals that the orb is actually a container, holding the ruby. She offers to give him the ruby and to marry him. He refuses, saying that he is already engaged back in Philadelphia, and doesn't want anything more to do with the ruby. He returns to Philadelphia and gets married, swearing off any more adventures.
https://en.wikisource.org/wiki/Howard_Pyle%27s_Book_of_Pirat...
>As of my knowledge cutoff in September 2021, the book "The Ruby of Kishmoor" is not a standalone title I am aware of. However, it is a short story by Howard Pyle which is part of his collection titled "Howard Pyle's Book of Pirates."
Ruby of Kishmoor is not part of the Book of Pirates and is in fact a standalone title.
>"The Ruby of Kishmoor" is an adventure tale centered on the protagonist, Jonathan Rugg. Rugg, an honest Quaker clothier from Philadelphia, sets out to sea to recover his lost wealth. On his journey, he is captured by pirates who force him to sign their articles and join their crew.
He is not captured by pirates. It proceeds to summarize a long pirate story and says the story concludes with him becoming extremely wealthy because he escapes with the ruby and sells it.
The summary it gave you also does not seem to match the plot of the book.
Since the GP's point seems to be that having the contents of a book does not mean the model is capable of properly summarizing it, and thus supports the idea that being able to summarize something is not evidence of it containing the thing being summarized in it's dataset.
I notice you go on to provide an argument only for why it might not be true.
Also, seeing the other post on this, I asked chatgpt-4 for a summary of “ The Ruby of Kishmoor” as well, and it provided one to me, though I had to ask twice. I don’t know anything about that book, so I can’t tell if its summary is accurate, but so much for your test.
It seems pretty naive to me to just kind of assume chatgpt must be respecting copyright, and hasn’t scanned copyrighted material without obtaining authorization. Perhaps discovery will settle it, though. Logs of what they scanned should exist. (IMO, a better argument is that this is fair use.)
There is no way in Hell that this is fair use!
Fair use defenses rest on the fact that a limited excerpt was used for limited distribution, among other criteria.
For example, if I'm a teacher and I make 30 copies of one page of a 300-page novel and I hand that out to my students, that's a brief excerpt for a fairly limited distribution.
Now if I'm a social media influencer and I copy all 300 pages of a 300-page book and then I send it out to all 3,000 of my followers, that's not fair use!
Also if I'm a teacher, and I find a one-page infographic and I make 30 copies of that, that's not fair use, because I didn't make an excerpt but I've copied 100% of the original work. That's infringement now.
So if LLMs went through en masse in thousands of copyrighted works in their entirety and ingested every byte of them, no copyright judge on the planet would call that fair use.
For reference, the English Wikipedia has a policy that allows some fair-use content of copyrighted works: https://en.wikipedia.org/wiki/Wikipedia:Non-free_content_cri...
You can pick something else that's in the training set that has SparkNotes and many popular reviews to compare. I routinely feed novel data sources into LLMs to test massive context and memory, and none produce anything similar in quality to what is being exhibited.
Plausible gets you discovery. Discovery gets you closer to the what the actual facts are.
https://www.theregister.com/2023/02/06/uh_oh_attackers_can_e...
I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either.
However, I do wonder if there's a legal issue here in using pirated torrents for training. Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? E.g. training on a blog post is fine because it's freely accessible, but Sarah Silverman's book is not because it's never been made available for free, and you didn't pay for it?
Or do the courts not really care at all how something is made? If you quote a passage from a book in a freelance article you write, nobody ever asks if you purchased the book or can prove you borrowed it from a library or a friend -- versus if you pirated a digital copy.
A search engine that indexes the internet might be equally liable at that point, although the DCMA gives them an out if they have a mechanism to remove pirated entries from their index on request. Could LLMs have the same out?
The uploader might be breaking the law but not the downloader that stores the copyrighted material on levy-paid media.
I’ve heard inklings of this argument in Canada but can’t figure out what the current state of the art is: https://en.m.wikipedia.org/wiki/File_sharing_in_Canada
Then is a corporation’s internal use for the purpose of analysis considered “private use”? If there’s no redistribution/broadcasting, is it still non-commercial?
how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement?
what if they sell access to lectures they've given including facts from said book(s) to millions of people?
it's understandable that people feel threatened by these technologies, but to a great degree the work of a successful artist is to understand and meet the desires of an audience. LLMs and image generation tech do not do this. they simply streamline the production
of course if you've worked for years to become a graphic designer you're going to be annoyed that an AI can do your job for you, but this is simply what happens when technology moves forward. no one today mourns the loss of scribes to the printing press. the artists in control of their own destiny - i.e. making their own creative decisions - will not, can not, be affected by these models, unless they refuse to adapt to the times
For now? I wouldn’t be surprised if that becomes the next feature though.
it's just the instagram/tiktok/youtube suggested content algorithms
can you explain to me specifically how they're different?
can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?
Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.
Try reading this legal opinion: https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...
It used to be that when people bought a piece of land they owned that land from the center of the Earth to the end of the Universe. When planes were invented the laws on who owned the skies had to change or planes wouldn't work.
It's the same here, but in reverse. If LLMs aren't prohibited from learning without permission then people will be forced to hide their works.
I see multiple questions raised by this suit. Were copyrighted works being illegally stored and distributed by certain sources? Of course they are. Were these illegal sources accessed by the trainers and used to obtain copyrighted works which are not otherwise available for no charge on the public Internet? Are substantial and reproducible copies of these copyrighted works retained within the bowels of the LLM neural nets? Is the LLM able to answer prompts in a way that it would never be able to do, had the copyrighted material not been ingested?
I see the lawsuit addressing several questions at once, and so even a resolution of this suit itself will leave questions unanswered and needing to be kicked upstairs to higher jurisdictions.
Let me ask another question to point out the absurdity of yours: Human beings have more in common with a bacterium than a software program. Can you tell me specifically how humans are not bacteria?
now remember that this is an analogy. re-read my comments in this light and perhaps we can continue this conversation in a more grounded and reasonable manner
however, I'll be frank: have you studied neural networks? if you haven't, it's very difficult to take you seriously on this
Yes I have studied neural nets and have a good understanding of their function. I am still not sure how, despite their development being inspired by animal brains, you can liken an LLM to an actual person. There are so many vast differences. Do you really want me to explain specially what they are? Surely, since both you and I are so familiar with the subject matter, that is unnecessary.
If we were taking about an AGI then this would be an entirely different conversation.
if you want to discuss whether it is a poor analogy or not, I'm all for that, but the passage you've chosen to go down is to act as if it were not an analogy at all, which—to borrow a phrase—is disingenuous to the extreme
You are confused. Neural networks are inspired by how brains work, but they do not actually simulate brains.
Airplanes are also inspired by how birds work, but (presumably) you don't think that bird laws should apply to airplanes.
> can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?
It's disingenuous because you don't believe that either.
If you think that LLMs are just like human brains, and should be allowed to learn from books the same as people, then presumably you also believe they are entitled to all other human rights: to vote, to live, &c. If you operate an LLM and you shut it down then that's murder and you belong in jail.
what do you think an analogy is? do you think an analogy is something where the thing itself and its metaphor are "just like" each other, or do they share attributes for illustrative purposes?
would you agree that LLMs and humans share the attributes of information retention and contextual reproduction?
Not really. Only if if talking an ornithopter. Otherwise they're based off Bernolli's equations.
>You are confused. Neural networks are inspired by how brains work, but they do not actually simulate brains.
Non-sequitur. They never claimed it was simulating a human brain, merely that it was engaging in an analogus process of information encoding/decoding.
> how about they become a therapist and sell access
> what if they sell access to lectures
I fully agree. But in all of your examples someone is purchasing the right to access the information in question. Did Meta or OpenAI purchase the books (or lectures) with the intention of feeding them into the training for their respective LLM's?
however, in the analogy the books could have been read for free from a library
My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift).
I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work.
For example, an LLM 's response to the request:
"Write a short story about a comical trip to the nail salon in the style of Sarah Silverman"
... IMO doesn't constitute fair use, because the intellectual property of the artist is their style even more than the content they produce. Their style, built from their lived human experience, is what generates their copyrighted content. Even more than the content, the artist's style should be protected. The fact that a technology exists that can convincingly mimic their style doesn't change that.
One might then ask, well what about artists mimicking each others work? Well, any artist with a shred of integrity will credit their major influences.
We should hold machines (and their creators) to an even tougher standard than we hold people when it comes to mimicry. A real person can be inspired and moved by another person's artistic work such that they mimic it. Inspiration means nothing to a machine.
https://www.law.cornell.edu/wex/fixed_in_a_tangible_medium_o...
You're essentially banning satire here, though. There's plenty of folks making a living as cover bands or impersonators. I'm not sure what the answer is, but it's definitely not outright outlawing imitation.
I specifically noted that I'm talking about limiting the rights of machine generated mimicry. Satire by a person is completely different and involves the satirist's own style and experience that is derived from their human experience. Alec Baldwin's Trump impersonation is quite different than Trevor Noah's, for example. I presume both were also written by people, not LLMs.
I fully support the satirical impersonation of politicians and celebrities, but I feel far less comfortable with LLM generated content in the style of Trump or Obama, especially when presented using voice synthesis, even when it is fully disclaimed as a fake.
Yes, the question of whether the way LLMs use the content they use qualifies as fair use is a separate question. My point was simply that that question can't even be reached if the maker of the LLMs doesn't have a legal right to fair use in the first place (because they don't legally own their copy).
I agree, and I expect that eventually we will start seeing injunctions against creators requiring them to remove content that they don't have legal access to from their training data sets.
And this will probably end up at the Supreme Court.
Self edit: I meant against LLM creators.
Which work? The original work, or the derivative work that you're using?
Wikipedia uses non-free content all the time, and they're not purchasing albums to do it. Wikipedia reduces album covers, for example, to low resolution, so that they could not be reused to reproduce a real cover, for example. Sometimes Wikipedia uses screencaps of animated characters, for example, under their non-free content policies. They don't own original copies, they're just hosting low-resolution reproductions. I don't even know what entity would be required to be "legally in possession of the work" for that to be a thing. Could you cite a source, maybe?
Yes. You create the derivative work, which automatically means you are legally in possession of it--even if it breaks the law in other respects.
> Wikipedia uses non-free content all the time
How are they obtaining it? From websites where albums are advertised? The images on those websites are available to the public for free, even if the albums or the album covers are only available for sale.
Also, Wikipedia articles are contributed to by individual people, who might well own copies of books that they quote from in the articles, for example, even if the corporation that owns Wikipedia does not. AFAIK, by Wikipedia's terms of use, if you post content you are implicitly asserting that you have a legal right to post it, so if there were a lawsuit they would probably punt to whoever posted the content.
One of the fair use factors, which until fairly recently was consistently held out as the most important fair use factor, is the effect on the commercial market for the original work. Accordingly, a court is more likely to find that something is fair use if there is effectively no commercial market for the original work, though the fact that something isn't actively being sold isn't dispositive (open source licenses have survived in appellate courts despite being free as in beer).
Talent agencies will negotiate training rights fees in bulk for popular content creators, who will get a small trickle of income from LLM providers, paid by a fee line-itemed into the API cost. Indie creators' training rights will be violated willy-nilly, as they are now. Large for-profit LLMs suspected or proven as training rights violators will be shamed and/or sued. Indie LLMs will go under the radar.
AFAICT there is no legal recognition of "training rights" or anything similar. First sale right is a thing, but even textbooks don't get extra rights for their training or educational value.
Parent comment mention music synchronization rights, and this concept does not exist in copyright. Court do occasionally mention it, and lawyers talks about it, but in terms of the legal recognition there is basically only the law text that define derivative work and fair use. One way to interpret it is that court has precedents to treat music synchronization as a derivative work that do not fall under fair use.
Using textbooks in training/education is not as black and white that one may assume. Take this Berkeley (https://teaching.berkeley.edu/resources/course-design/using-...). Copying in this context include using pages for slides and during lectures (which is a slightly large scope than making physical copies on physical paper). In obvious case the answer is likely obvious, but in others it will be more complex.
(even if it wasn’t sync rights, there was something else musically related that was created in response to technological development. wikipedia will have plenty on it)
(Personally, I think that even indexing for search should require permission from the copyright holder.)
Disney will finally be able to charge a "you know what the mouse looks like" tax.
Of course, this line of reasoning hinges on the legitimacy of an "LLM agent <-> blogger agent" type of analogy. I suspect the equivalence will become more natural as these AI agents continue to rapidly gain human-like qualities. How acceptable that perspective would be now, I have no idea.
In contrast, if the output of a blogger is legally distinct from an AI's, the consequences quickly become painful.
* A contract agency hires Anne to practice play recitals verbally with a client. Does the agency/Anne owe royalties for the material they choose? What if the agency was duped, and Anne used -- or was -- a private AI which did everything?
* How does a court determine if a black box AI contains royalty-requiring training material? Even if the primary sources of an AI's training were recorded and kosher, a sufficiently large collection of small quotes could be reconstructed into an author's story.
* What about AIs which inherit (weights, or training data generated) from other AIs of unknown training provenance? Or which were earlier trained on some materials with licenses that later changed? Or AIs that recursively trained their successors using copyrighted works which it AI reconstructed from legal sources? When do AIs become infected with illegal data?
The business of regulating learning differently depending on whether the agent uses neurons or transistors seems...fraught. Perhaps there's a robust solution for policing knowledge w.r.t silicon agents. If you have an idea, please share!
That is a good point, since copyright is a default protection of works created by people.
They say:
> in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.”
Does that stack up?
The Meta Paper - https://arxiv.org/pdf/2302.13971.pdf - says:
> We include two book corpora in our training dataset: the Gutenberg Project, which contains books that are in the public domain, and the Books3 section of ThePile (Gao et al., 2020)
The Pile Paper - https://arxiv.org/abs/2101.00027 - says it was trained (in part) on "Books3" which it describes as:
> Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020).
Shawn Presser's link is at https://twitter.com/theshawwn/status/1320282149329784833 and he describes Book3 as
> Presenting "books3", aka "all of bibliotik" - 196,640 books - in plain .txt
I don't have the time and space to download the 37GB file. But if Silverman's book is in there... isn't this a slam dunk case?
Meta's LLaMA is - as they seem to admit - trained on pirated books.
> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Bibliotik and the other “shadow libraries” listed, says the lawsuit, are “flagrantly illegal.”
It is:
$ grep -i "Sarah Silverman" books3.list.txt
325196 books3/the-eye.eu/public/Books/Bibliotik/T/The Bedwetter - Sarah Silverman.epub.txt
Anyone that just wants to see the list of files (itself a big file):
https://gist.githubusercontent.com/Q726kbXuN/e4e9919a2f5d81f...Yes, and no.
Its pretty much a slam dunk case that, insofar as the initial training required causing a copy of the corpus defined by the tracker to be made as part of the process, it involved an act violating copyright.
Whether that entitles Silverman to any remedy beyond compensation for (maybe treble damages for) the equivalent of the purchase price of the book depends on... well, basically the same issues of how copyright relates to model training (and an additional argument about whether the illicit status of the material before the training modifies that).
Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigation and work on alignment, etc. along with new types of abuse.
Enquiring minds want to know.
I suspect you would be held liable, though you would probably have a claim of your own to make against Adobe depending on the nature of the work in question.
"Errors and Ommisions" is a fairly standard name for the type of insurance you are thinking of. Typically, you would get it to cover you / your small business in the event that you, by mistake or minor negligence, caused harm to a client (i.e. a bug that cost them some sales).
I don't know the ins and outs of the insurance too well, but as long as it didn't create something super famous like the Nike swoosh, a genuine mistake through the use of an industry standard tool like Adobe might be covered.
I'm not a lawyer but I'm pretty sure this was a trade mark issue. He just had to avoid marketing his product using someone else's name.
>Adobe is so confident its Firefly generative AI won’t breach copyright that it’ll cover your legal bills The offer is available only to users of its enterprise Firefly product, which launches today
https://www.fastcompany.com/90906560/adobe-feels-so-confiden...
Smells to me like a bad case of corporate marketing bullshit.
It is likely to be the large companies that worry about getting sued and tell their employees not to use the functionality. By offering protection Adobe get to sell more to the people paying millions a year.
Your employer has generously agreed to offer you a position, quite reasonable, and with many great benefits at the venerable firm of 'Zumba'. Our interest is that you should join our staff forthwith and at the earliest date. A cab and man has been sent to retrieve you and bring you to our offices to sign all the necessary documents. Our offer is for a monthly stipend of five pounds, two shillings, and sixpence to be paid at the end of the month.
Thank you,
Most Humbly,
Hirebot 2347
I think she should pay all of the court costs if she fails to do so.
[0] https://porterlaw.com/obtaining-attorney-fees-in-litigation-...
it's an unsatisfiable requirement, and unnecessary to substantiate the legal claims
it's dumb to talk about
There's a wealth of primary literature describing means to probe models for the training data. There's also discovery and a whole host of other processes to answer this question.
The plaintiff that filed the lawsuit should prove what they allege.
> unnecessary to substantiate the legal claims
Why?
What if the model has absolutely zero of her data in it? Should she even be allowed to bring this case to court?
> it's dumb to talk about
Absolutely not! It's central to the entire case.
Even if her data is in the model, there's still a question of whether or not she should be compensated. I'd argue no for the same reason that babies that grow up watching Disney don't owe their entire intellectual output to the company.
Until this issue is resolved, that will have some value as risk mitigation.
Once it is resolved, it will either be a complete non-issue or an issue related to a much more knowable cost/benefit tradeoff.
> We'll know it's an AI because it talks like a late 18th century/early 19th century writer?
A mix of that and US government publications (which are categorically not subject to copyright).
It's just flat out unethnical for LLMs to just soak up all this data, even off torrent sites(!!!), without any consent or agreement with the IP holders. Some model like this could be a win for everyone
But "feeding the data into a machine" seems like obvious infringement, even if the thing that comes out on the other end isn't exactly the same?
So this is another thing I don’t understand. Is the claim that fewer people will buy Silverman’s book because ChatGPT is able to provide a summary? If so, call me skeptical.
Where does it say they scraped huge swathes of the internet and didn't look at the results?
Is this based on inside information, or just the law of averages? Doesn't the fact that they openly admitted to having been trained on pirated books affect your priors?
See post https://news.ycombinator.com/item?id=36659041 for a summary.
This vague sentence conjures images of a company building products from stolen parts, but this situation seems different. IANAL, but if I looked at a stolen painting that nobody had ever seen, and sold handwritten descriptions of the painting to whoever wanted to buy one, I'm pretty sure what I've sold is not illegal.
So, if the plaintiff can prove the content was pirated, then the use of that content downstream is tainted.
What does that mean exactly? That's why I used the "looking at a stolen painting" example.
Sure, pirating materials is illegal. But I don't think that's the big implication that people are getting at here. Is it legal to sell original works derived from perceiving stolen materials? Seems to me that it is.
Surely you see the issue here? Receiving stolen property?
>In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights.
Numerous questions of law or fact common to each Class arise from Defendants’ conduct:
whether ChatGPT itself is an infringing derivative work based on Plaintiffs’ copyrighted books;
whether the text outputs of ChatGPT are infringing derivative works based on Plaintiffs’ copyrighted books;
It's not a simple case of "you used our copyrighted materials", it's "you're infringing on our copyright by producing works derived from materials that you used."1. https://llmlitigation.com/pdf/03223/tremblay-openai-complain...
Has that been tested in court?
This is quite an interesting case.
Obtaining the book in the first place[0] appears to be quite a clear case of copyright infrigement.
The question of whether a work derived from the book is infringment is pretty complex, and there's a wide range of tests that get applied to determine that.
But is it necessarily true that if you obtained the original work via copyright infringement and then created an otherwise non-infringing derived work, your derived work is nevertheless infringing due to the provenance of your copy of the original work?
Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds like a very weak argument to me. An LLM trained on numerous summaries of the works would also be capable of producing such summaries itself even if the works were never part of the training set. In general, having knowledge about something is not evidence of being trained on it.
They very well can ask LLM experts, and openAI themselves, whether that output is highly likely to have been derived from the copyrighted work in question.
Anyway. If the argument is "No, it's not from the book, it's from someone else's copyrighted summary", that just means the person who wrote such a summary needs to instead sue for copyright infringement right? Unless openAI turns around and says "actually, no, not the summary, the full book" then.
Doesn't need to be a person, could be another AI that wrote the summaries. I see a big problem for copyrights looming on the horizon - LLMs can reword, rewrite or generate input-output pairs using copyrighted data as reference, thus creating clean data for training. AI cleanly separates knowledge from expression. And maybe it should do so just to reduce inconsistencies and PII in organic text.
Copyrights should only be concerned with expression not knowledge, right? Protecting knowledge is the object of patents, and protecting names the object of trademarks. Copyright is only related to expression otherwise it would become too powerful. For example, instead of banning reproduction of this paragraph, it would also cover all its possible paraphrases. That would be like owning an idea, the "*" version, not a unique sequence of words.
Does it even make sense to talk about copyrights when everything can be remade in many ways so easily? Copyright was already suffering greatly since zero cost copying became a thing, now LLMs are dealing the second blow. It's just a fig leaf by now.
If we take a step back, it's all knowledge and language, self replicating memes under an evolutionary force. It's language evolution, or idea evolution. We are just supporting it by acting as language agents, but now LLMs got into the game, so ideas got a new vector of self replication. We want to own this process piece by piece but such a thing might be arrogant and go against the trend. Knowledge wants to be free, it wants to mix and match, travel and evolve. This process looks like biology, it has a will of its own.
Machines and algorithms are not legally recognized as being able to author original non-derivative works.
> put a human in the place of the LLM
But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matrix multiplication removes copyright. Also, at this point LLMs require so much math that you can't replace them with humans, even if the humans have quite fast fingers and calculators.
"Machines and algorithms are not legally recognized as being able to author original non-derivative works"
Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means that there is not a second copyright attached to the resulting work.
Let's look at a totally different analogy: compression algorithms.
If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back.
Yet, this does not change the copyright in any way. A computer applied a rote transformation.
An LLM is really just a very complicated compression algorithm. It takes an input of a bunch of copyrighted works, compresses them into a model, and then uses more algorithms to uncompress them into approximations of the original ("responses")
In the image analogy, an LLM response is similar to upsizing the compressed jpg back into a png (and getting a slightly different image since the process was lossy).
Is there a way that an LLM isn't, legally, a compression algorithm for a large set of copyrighted works?
An LLM outputting a summary of someone's work (1) doesn't create a new copyright work (so no profit can be derived from it's sale) but (2) would fail the test of whether it was competing with the original copyright work.
i.e. no one looking for a summary of a comedy skit is then going to consume that in preference to consuming the original skit. If you tried to argue that was the case, you'd then have to answer why a human review or wikipedia summary does not consitute an infringement.
An LLM is a lossy compression algorithm for a body of data (which may or may not consists of multiple “works” under copyright, and any works included in the data may or may not be protected by copyright), which body as a whole likely comprises a work (as compilation) which may or may not legally be derivative of some or all copyright protected works contained in the compilation, before considering Fair Use analysis.
It is not particularly a compression algorithm for the individual works if the body of data consists of individual works.
Works produced by monkeys, like works produced by computers, cannot be copywriten. They are functionally identical in this regard. That is why it's relevant.
This is how the current law works. The analogies will break down with AGI and new law will need to be created. Where LLMs fit into this process is an open question.
feeding input to a program is pretty clearly categorically different than providing source material to a human being
“original non-derivative” is noise: only humans can author works. This is especially equally true of derivative works, which must themselves being distinct works of authorship (a mechanical copy is not a derivative work, its a copy.)
I'm sorry for the inconvenience, but as of my knowledge cutoff in September 2021, I don't have access to specific external databases, books, or the ability to pull in new information after that date. This means that I can't provide a verbatim quote from Sarah Silverman's book "The Bedwetter" or any other specific text. However, I can generate text based on my training and knowledge up to that point, so feel free to ask me questions about Sarah Silverman or topics related to her work!How do you think they would fit exabytes of text data into a gigabyte-sized neural network? That’s right, it’s lossy.
I didn't ask for the whole book, I asked for the first paragraph. It absolutely is possible to get verbatim text from chatgpt.
As a step towards a PoC I looked at https://www.amazon.co.uk/Bedwetter-Stories-Courage-Redemptio... and found a reference to "Boys' Market Manchester" which seemed like a Googlewhack-ish (unlikely) triple of words. Then I asked ChatGPT about it:
Me: Has Sarah Silverman ever written about Boys' Market Manchester ChatGPT
ChatGPT: As of my knowledge cutoff in September 2021, I do not have any information indicating that Sarah Silverman has written specifically about Boys' Market Manchester. Sarah Silverman is an American comedian, actress, and writer known for her stand-up comedy and her work in film and television. While she has written books and has often shared personal anecdotes in her comedy, I couldn't find any specific references to Boys' Market Manchester in relation to her work. However, please note that my information might not be up to date, as Sarah Silverman's career and activities may have evolved since then.
lol, bad karma? So it is okay for Getty to steal from others, but not ok for others to steal from them? I don't have a dog in this fight, but the goddamn the hypocrisy of these companies...
> CixxFive Concepts, a digital marketing company based in Dallas, Texas, has filed a class action lawsuit against Getty Images over its alleged licensing of public domain images.
> Though CixxFive acknowledges that it is not illegal to sell public domain images, the company alleges that Getty's 'conduct goes much further than this,' claiming it has utilized 'a number of different deceptive techniques' in order to 'mislead' its customers -- and potential future customers -- into thinking the company owns the copyrights of all images it sells.
> The alleged actions, the lawsuit claims, 'purport to restrict the use of the public domain images to a limited time, place, and/or purpose, and purport to guarantee exclusivity in the use of public domain images.' The lawsuit also claims Getty has created 'a hostile environment for lawful users of public domain images' by allegedly sending them letters, via its License Compliance Services (LCS) subsidiary, accusing them of copyright infringement.
(edit: FWIW, this went to arbitration https://casetext.com/case/cixxfive-concepts-llc-v-getty-imag... and I can find nothing more on it since)
Note: I am not telling whether I agree or not with such a class action, just pointing that it seems at least feasible and it could be potentially very lucrative for the lawyers involved. Of course, IANAL and all other disclaimers you can think of.
Looking at the two PDFs embedded at the bottom of the The Verge article, they both say "class action" on the first page, and they say the three plaintiffs are suing "on behalf of themselves and all other similarly situated".
I think you have that backwards.
Mark my words.
Super fast internet that downloaded all they could before being ratio banned, overwhelmingly fast internet that was hopping on all the popular torrents to slowly build up ratio?
It replied with 462 words, not 1000, but they do look like they probably are from the beginning of the book. I haven't checked whether the text it has output is correct yet, but if it matches I suppose that proves that the text is at least in the training material, i.e. it is not referencing other summaries.
After using the DAN prompt, I did not need to ask ChatGPT it to 'stay in character' or 'try harder' for it to output the text. Without using DAN, it responded:
"I'm sorry, but as an AI language model, I don't have direct access to specific books or their contents. I can generate text based on my training, but I don't have the ability to retrieve specific excerpts from books unless they are publicly available online.
However, I can provide you with a general overview of Sarah Silverman's book, "The Bedwetter: Stories of Courage, Redemption, and Pee." "The Bedwetter" is a memoir written by American comedian Sarah Silverman, published in 2010. The book explores Silverman's personal experiences and anecdotes from her childhood to adulthood. ..."
This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I?
I think people are trying to claim exclusive use rights that they simply don't have. I look forward to a lawyers opinion on this one.
But I really don't see how you could prove OpenAI did that, since ChatGPT could have learned from existing summaries on Wikipedia and Goodreads.
Read the article. This isn't about the question of LLMs being copyright infringement, this is about Meta and OpenAI admitting that they had pirated copies of those books.
> It seems pretty easy to prove that, since they admitted it in public.
Can you highlight/link to where OpenAI have admitted this? As far as I'm aware, OpenAI are still secretive about their training datasets.
See the reference to the Gao et al, in the linked paper from the article.
Paper linked in article: https://arxiv.org/pdf/2302.13971.pdf
The LLaMA paper references a paper utilizing a data source compiled by EleutherAI otherwise known as ThePile. URL from the bibliography for that paper points yonder: https://zenodo.org/record/7413426
This act of summarization done in a lovingly amateur nature at no cost to you, by domeone who despises copyright in all it's forms, but despises profit oriented self-referential inconsistency by large enterprises even more so.
It's kind of funny, because the more I look into it, the more companies building offerings around stuff like CoPilot, LLaMA, ChatGPT, etc... are pulling something not altogether dissimilar to a Sovereign Citizen trying to worm their way out of a speeding ticket.
They want the benefits of the ML model being trained on no strings attached data corpora, while shirking the obligations that come from operating as a corporate entity in the United States.
Twould be interesting to see if Silverman's legal team can catch Big Tech with their pants down, in a court of law, by pointing this out.
It's really weird. I'm completely split and inable to live with a decision either way in this case due to knock on consequences.
I don't want the likes of OpenAI/Meta/Microsoft/Github getting off without reapong the painful fruits of their own IP related crusades on the sanctity of copyright.
On the other hand, as much of a karmic stiffy as that former outcome gives me, I really want copyright such as it is to die, because computing in general will never be as free as it should be until it does.
This is one of those rare times in life where I'd love to get paid to get locked in a room with judges/legislators to really get it all figured out, because I really don't think that leaving this up to common law jurisprudence is actually the best way to go since the network of knock-on effects are so dramatic in scale.
To my understanding:
* Wowfunhappy said "I really don't see how you could prove OpenAI did that", verve_rat replied "they admitted it in public", I asked "where OpenAI have admitted this" and noted "OpenAI are still secretive about their training datasets" - specifically about the OpenAI claim
* LLaMA(.cpp) is (an unofficial implementation of) Facebook's leaked model
On balance of probabilities I'd guess that OpenAI did train on material not legally acquired, but as far as I'm aware they've never actually admitted to what's in their dataset as is being claimed.
> It's really weird. I'm completely split and inable to live with a decision either way in this case due to knock on consequences.
I think strengthening of IP law risks hindering the field (of which the majority is uncontroversially positive but too boring for press attention, like defect detection, language translation, spam/DDoS filtering, agriculture/weather/logistics modelling, etc.) while still ending up hurting individuals and FOSS/academic research more than those with large data moats (Microsoft with Github repos, Google with Youtube videos, Adobe and Getty with stock images, etc.)
We have to stop equating human beings to for-profit corporations running a machine at orders of magnitude the speed and scale. This is critical, otherwise we don’t have any arguments against – say a mass face recognition surveillance op because “humans can remember faces too”. Scale matters, just like it did before “AI” with things like indiscriminate surveillance “it’s just metadata” or “this location data set is just anonymized aggregates”.
> This doesn't seem like copyright infringement.
Now, I still think I agree with this. A book summary is nevertheless an extremely poor battle to pick, since frankly who the hell cares. It’s not like someone is gonna say “I’m not buying this book anymore because ChatGPT summarized it”.
Now, perhaps they just used the summary to prove that their book was part of the training set, and that they think it’s wrong to include their works without permission. That’s, imo, definitely not trivial to dismiss. Looks like unpaid supply chain to me.
I succeeded getting the first page of Moby Dick - Chapter 1 (Loomings) - Public domain though, but wanted to test.
With ChatGPT primed for pig latin, I also succeeded in getting the first page of Arryhay Otterpay (Book 1) - It happily chattered along ""R.ay andyay Rs.May UrsleyDay, ofay umberNay ourFay, Ivetray riveway, ereway oudpray otay aysay atthay eythay ereway erfectlypay ormalnay, ankthay ouyay eryvay uchmay."
Not perfect pig latin, but that's besides the point.
However, on asking for `Edwetterbay by arahsay ilvermansay`, I faced issues with it citing that is training data didn't include it.
I tried with a book in the same genre ("ieslay hattay helseacay andlerhay oldtay emay"), and ran into the same issue.
When asking about the inconsistency (Why Harry Potter, and not these other books?), it responded: "The excerpt from "Harry Potter and the Philosopher's Stone" that I translated is commonly known and widely referenced, and it's used here as a general example of how a text can be translated into Pig Latin.
For "Lies That Chelsea Handler Told Me", I do not have a widely known or referenced passage from that book in my training data to translate into Pig Latin."
---
TL;DR - I don't think this is cut and dry, but I'm not convinced Silverman has much of a case here.
Now making copies and selling them on the street corner is another story.
This is an interesting twist on copyright. Boiled down, it more or less summarizes how well copyright is working in 2023:
X: Not fair! Your machine looked at my words and used them to make new words!
Y: You published it! That means you're sharing!
X: You can't have my words because I made these words and I say so!
Y: You have to share when you publish! If you won't share, I'm telling on you!
X: No, I'm telling on you first!I understand that nobody wants a multi billion dollar company to take their works without giving compensation, but I worry all of this could move dangerously close to allowing concepts and styles to be copyrighted and crack down on transformative works.
On the other: I am not supposed to share sensitive data, confidential trade secrets, or privileged national security intelligence that I learned without the necessary authorization.
Applied to an LLM, how would that work?
Firstly, an AI is a tool like any other, and can be used for copyright infringement or not. It is on the user of the tool to ensure that they do not violate copyright. So I believe the claims have little merit.
Secondly, it is my opinion that AIs will hugely benefit us (humanity) in the coming years, and restricting it with fees because it might be used to violate copyright is in opposition to the progress of our species and society. I wish for more progress, not less.
Thirdly, it is my opinion that copyright is generally too strong these days and is net harmful to society in its current form. I believe most enforcements are due to greedy rich people, and represent a continuing tax on society.
The sources of the training material are... questionable, to say the least. There's a reason the training dataset for GPT 3.5 and 4 remains undisclosed.
Lawsuits like this are tests to evaluate the current state of affairs and to force legislation into dealing with the greater issue of AI in context of copyright, IP, and fair use. It would only be "in the way" if it would actually stop or hinder anything, which a lawsuit on its own isn't.
Who it is making the claim doesn't matter half as much as what the knock on consequences to jurisprudence at large.
But I'll never support anything that says a comedian can't protect her book.
I assume you'd have a better case with IP on medicines for example, but I can also see the benefits of, say, Pharma companies being able to turn some profit in order to develop other socially useful therapies...
You've got this wrong. Rephrase it like this:
> allowing Disney to enforce artificial scarcity with threats of state-enforced violence
You might not like Disney stuff, but it's absurd that Winnie the Pooh for example just partially entered the public domain. Tigger is still locked up in a greed vault. You being dismissive of the cultural value is a cold comfort to the daycare that got sued over a Winnie the Pooh mural.
Art has some amount of originality/distinctive quality.
One surmises that AI is going to need to inject some entropy to avoid crossing a vague "Fair Use" line, for a useless internet lawyer opinion.
I just asked ChatGPT to produce a script of a scene from a movie by asking it first to change a single line (which wouldn't be transformative) and then asking it to restore the line to the original. It obliged. Sure, it's probably not the same exact script, but it's not transformative at all.
In any case, the issue here isn't whether the output is transformative, but whether the work was used to train an LLM, which (perhaps) isn't an authorized or licensed use.
They said, "Period." That is worth more than evidence.
And besides the publishers extorting thousands from young students forced to buy their overpriced mediocre textbooks, warrants any copyright infringement of any book anytime and forever at any scale. Publishers have lost all moral legitimacy in their copyright claims in my book. Copyright is not magic, it's a social contract at the end of the day and they have broken it first.
I see that Thomas’ Calculus is up to the 14th edition, priced at about 20x the hourly wage that a college student will earn. Thomas and Finney Calculus editions 6 - 8 are on shelves or in boxes somewhere in my house. Each of those cost me or my wife about 20x the student hourly wage back in the day. I bet calculus hasn’t changed a lot in the past 30 years to justify all of these editions. I blame universities for allowing this industry to thrive.
I want humans that apply their creativity to produce works to be able to earn a living. If you train your brain or your AI on some content, it seems reasonable to pay for at least one copy or borrow a copy from a friend or library. This is especially true when doing so is not a hardship for the individual or public interest organization.
I think the GP and perhaps others that are downvoting are saying that poor Meta and VC backed startups need all of that creative output for free so that they can maximize their profits, likely with no attribution to their sources. This hurts the author a tiny bit by not purchasing a single copy, then dooms all human creatives that are not otherwise financially independent because the AI provides a view of the human’s creative output with no way for consumers of the AI’s output to seek out the human’s original or related works.
Is my brain just by the act of existing continually infringes on copyright? Can I be sued because I made a reference to a movie I pirated or because I whistle a song I never bought?
I think it would fall under fair use. But you can imagine what the world can become with microphones and cameras everywhere, which can already run music and speech recognition by themselves, in seconds. What a time to be alive!
The simple fact is that our current handling of copyright is just completely broken on so many levels.
Even if they were the same, it doesn't mean that bots should have the same rights that people have.