NY Times copyright suit wants OpenAI to delete all GPT instances
arstechnica.com
arstechnica.com
Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets.
Is that fair use? IANAL, but doesn't sound like it. Typically I can't take a personal "tier" of a product and charge 3rd parties for derivatives of it. Say like VS Code.
A sibling comment mentions search engines. I think there's a big difference. A search engine doesn't replace the source, not at all. Rather it points me at it, and offers me the opportunity to pay for the article. Whereas either this or an LLM uses NYT content as an alternative to actually paying for an NYT subscription.
But then what do I know...
>Is that fair use? IANAL, but doesn't sound like it.
If you pay someone to do the summarisation for you, then you publish the content and charge a fee for it, you're the one liable, not the person you paid to summarise it for you. Similarly if you ask GPT to do it for you, then publish it, you're liable for what you publish; GPT is just a summarisation tool.
At some level it becomes a subversion of NYTs fees. First, say I subscribe and simply host the articles verbatim, for a fee. Clearly, that's not right.
Suppose I change some spelling or word order, or use a synonym or two. That's still not ok.
And if I substantially paraphrase the articles? I guess this is the relevant case. This is kind of what LLMs do. And also feels like not fair use.
That's not what OpenAI is doing; it's not selling summarised articles as a service. Your example is a false equivalence.
>This is kind of what LLMs do. And also feels like not fair use
An LLM doesn't do this unless you ask it to. And if you then take that output and publish it as your own, you're breaching the copyright, not OpenAI.
In this case, OpenAI is violating copyright by modifying, reproducing and distributing copyrighted content to its customer.
I read a NYT article, then summarize it into a link title for reddit. Reddit then republishes the summary to all of its users.
So, if the summaries are derived works and not covered by fair use, then both you and the summarizee are separately breaking the NYT's copyrights. Otherwise, if this is covered by fair use, then you are both in the clear.
Finally, GPT is not "a summarization tool" in this case. If you provide a copy of a NYT article as a prompt and then ask for summarization, then yes, it is clear that GPT is not doing anything wrong, even if it spits out the exact same text. But if you simply ask for a summary of a specific article by, say, just name and date, and you get a copy of it, it's clear that GPT is storing the original data in some way, and thus it has copied the NYT's protected works without permission.
In this particular case they were using it via Bing, which actively did a HTTP request to the particular article to extract the content. So GPT hadn't memorised it verbatim, instead it fetched it, much like a human using a search engine would.
Additionally, even the use through Copilot is very debatable. They are not returning the NYT link, which requires a subscription, they are returning the contents of it even to non-subscribers. And they are doing this in a commercial product, not a non profit like the Internet Archive, which has some arguments for fair use.
Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.
But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn".
Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying property for some amount of time also has squatter's rights, even though you could easily and affordably automate and scale the deployment of drones to live and hide away on properties long enough to attain rights regarding properties all over the country.
I await the HN ban with fear..
[1] I'm not even doing referencing - so I am surely an LLM.
but, more importantly, OpenAI can also be sued for tortious interference? (basically the civil equivalent of accessory)
That's function of the legal system, not of the technology. If tomorrow someone made a perfect dolphin-Esperanto translator and proved Dolphins were as smart as humans, you still can't sue a dolphin until the legal system says so.
Not exactly, no, but the 'neurons that fire together wire together' way of learning has a pretty similar effect.
> LLMs "learn" in the same way that Excel sheets "learn".
I've never seen an excel sheet do anything like backpropagation.
Not strictly in the sense you mentioned (assuming that you mean "by themselves") but people may find [1] and [2] interesting.
[1] https://pub.towardsai.net/building-a-neural-network-with-bac...
[2] https://towardsdatascience.com/demystifying-feed-forward-and...
Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together.
And ultimately, at the deepest level we can analyse, our brains' atoms are doing quantum field diffusion equations, which you can also do in an Excel spreadsheet, so that kind of reductionism doesn't help either.
> Humans are living beings with needs and rights. A person being able to legally squat in a home doesn't mean that a drone occupying property for some amount of time also has squatter's rights, even though you could easily and affordably automate and scale the deployment of drones to live and hide away on properties long enough to attain rights regarding properties all over the country.
Yes, but we can also do tissue cultures and crude bioprinting, so it's a very foreseeable future where exactly the same argument will also be true for living organisms rather than digital minds.
We need to figure out what the deeper rules are that lead to the status quo, not merely mimic the superficial result. The latter is how cargo cults function.
Sure, that's an interesting path of inquiry, and one should be free to understand themselves as being no different than a machine if they desire.
But the objective of laws is the benefit of (at least some) humans, not machines covered in lab grown tissue. The process of being human is a big part of what makes us human.
I think you're misapprehending — I mean an entity fully 3D printed out of tissue, no machinery (unless you're counting all biology as machinery, but I think you're not doing that).
I recon bio-printing is now where home computing was in the Apple 1 era, so this is a way off, but it's foreseeable.
> The process of being human is a big part of what makes us human.
Mmm. How much has that process that changed since the ancient world?
How do you recon that, Apple 1 was Turing complete. We haven't printed life, that would be a tremendous accomplishment.
I think we're closer to Edison inventing a lightbulb as a step to computers being possible. Printing a conscious thing, at all, would be like the transistor. An Apple 1 analogue wouldn't be likely because of the terrible ethics of a "shitty" printed human.
Sure we have, and in multiple different senses.
The ones which matters here are cell culture, which is nowhere near the fanciest bar that's been surpassed in this field, and tissue culture which is somewhat harder but the reason why I recon it's at the Apple 1 level is that a small number of experimentalists are messing around with it using expensive equipment that you can technically buy at home but you need to be well trained to actually use, for example:
https://youtu.be/Z_ZGq8Tah0k?si=u6bBatjuSWcyNYJ3
And, more broadly, there's bioprinting as a research etc. field:
https://en.wikipedia.org/wiki/3D_bioprinting
And here's a TED talk from ten years ago where they demoed an early research 3D printed kidney on stage:
No. That isn't printing life, that is taking already living cells, priming and transforming them into something useful. Regardless, I'd count it if we could make an entire living organism this way, but we cant. Creating a working organ is no doubt amazing, and proof that this technology is worth pursuing, but it isn't "printing life" any more than producing life saving drugs is.
In your example you are talking about being able to bioprint a person(they have to be a person to have that right) to squat a property. Bio printing an organ isn't an example of that, it's not even close. Saying that we are anywhere near being able to print a human to squat a property is pretty ridiculous.
Which is absolutely sufficient for the usage I described upthread. In fact, I'd go so far as to say it's mandatory for the point I was making, as — fun though bio-printed werewolves, dragons, and fae would be — my point only works if you get humans out of the process rather than some other species. A bioprinted horse is probably slightly harder than a bioprinted human, but the latter isn't getting any squatting rights.
I could've linked to work on synthetic genomes and nucleotides to give evidence for lower-level creation of live, but they don't matter for the same reason:
My point is that there's a pathway heading off into the distance, and somewhere in the distance but before the horizon can be found bio-printed humans with all the same moral issues we're now just beginning to taking seriously thanks to AI being conversational, and if we had something completely customised, that's cool and all, but it doesn't make anyone go "oh, they're people" the way a humanoid body with human DNA getting off a table saying "hello, nice to meet you" does.
> In your example you are talking about being able to bioprint a person(they have to be a person to have that right) to squat a property. Bio printing an organ isn't an example of that, it's not even close. Saying that we are anywhere near being able to print a human to squat a property is pretty ridiculous.
I wrote "an entity fully 3D printed out of tissue […] is a way off, but it's foreseeable" and compared bio-printing today to a nearly 50 year old computer, and one of my references was a link to a youtube channel where someone is attempting to do a small-scale prototype thing along these lines with a handful of organs made from mouse cells grown in his own lab (and mouse cells rather than human because of the disease risk not because something magic happens with human cells). You're mixing up what I think is foreseeable with what I say already exists, and using the nonexistence of what I think can be foreseen to argue against what does exist.
No! Hebbian learning is categorically NOT gradient based learning. Hebbian update rules are local and not the gradient of any function.
Cortical learning is so vastly different from how artificial neural networks “learn” they cannot even begin to be meaningfully compared mathematically. Hebbian learning is not optimization and backprop is not local learning.
Part of the problem of these discussions is a bunch of clueless people talking with authority.
I have to keep reminding myself that outside of my own speciality, ChatGPT knows more than me despite its weaknesses, so I bet ChatGPT knows more about Hebbian learning than I do.
I'll look into that more.
most of the world disagrees with this view, and that means they will create the AI that wins.
You misunderstood me. I was talking about something more fundamental.
Understanding is data compression. They are the same thing. Learning patterns, building mental models, creating abstractions, generalizing, gaining intuition/a feel for something - all the things humans engage in as part of learning and understanding the world - are all acts of lossy data compression.
Anyone got more details on this?
Superficially it sounds like total BS; a highly compressed zip file does not exhibit any characteristics of learning.
Algorithmically derived highly compressed video streams do not exhibit characteristics of learning.
?
I’ve vaguely heard the learning can be considered to exhibit the characteristics of compression in that understanding of content (eg. segmentation of video content resulting in more highly compressed videos) can lead to better compression schemes.
…but saying you can “do a with b” and “a and b are fundamentally the same thing” seems like a leap…?
It seems self evident you can have compression without comprehension.
An LLM has limited parameters. If an LLM had infinite parameters it could just memorize the results of every single addition question in existence and could not claim to have understood anything. Because it has finite parameters, if an LLM wants to get a lower loss on all addition questions, it needs to come up with a general algorithm to perform addition. Indeed, Neel Nanda trained a transformer to do addition mod 113 on relatively few examples, and it eventually learned some cursed Fourier transform mumbo jumbo to get 0 loss https://twitter.com/robertskmiles/status/1663534255249453056.
And the fact it has developed this "understanding" as an ability to learn a general pattern in the training data enables it to compress. I claim that the number of bits required to encode the general algorithm is fewer than the number of bits required to memorize every single example. If it weren't then the transformer would simply memorize every single example. But if it doesn't have space then it is forced to try to compress by developing a general model.
And the ability to compress enables you to construct a language model. Essentially, the more things compress, the higher the likelihood you assign them. Given a sequence of tokens say "the cat sat on the", we should expect "the cat sat on the mat" to compress into fewer bits than "the cat sat on the door". This is because the latter is far more common and intuitively more common sequences should compress more. You can then look at the number of bits used for every single choice of token following "the cat sat on the" and thus develop a probability distribution for the next token. The exact details of this I'm unclear on. https://www.hendrik-erz.de/post/why-gzip-just-beat-a-large-l... this gives a good summary.
I fundamentally disagree. That's not some established fact, just a narrative used by those who wish to plagiarize using "AI".
Our collective human limitations(physical, mental and temporal) are sort of invisible implicit rules that we all follow in one way or the other. If an entity is not bound by those rules then I don't see why that entity should be treated the same as a human.
Companies already make this differentiation.
For example take captcha and bot detection. Some of the heuristics are based on inherent human limitations like response time, click time, mouse acceleration etc.
I doubt youtube or any other streaming service will be happy if you want to stream all their videos to train a hypothetical human like AI(which views and prepares notes like a human) at a hugely accelerated speed compared to a regular human. You can guess how quickly they will cite fair usage policies.
What I want to say is there are fundamental differences between a human and an AI. So, we should not be quick to dismiss any concerns just because AI can "mimic" humans in certain areas.
Humans have rights, software tools don’t.
If you grant an LLM the full set of human rights, then it can consume information, regurgitate copyrighted works, and use it to generate money for itself. However, considering blatantly obvious theft as “homage” goes hand in hand with free will, agency, being in control of yourself, not being enslaved and abused, etc. Pondering various scenarios along those lines really gets to the heart of why an LLM is so very much not a human, and how subjecting it to the same treatment as humans is a ridiculous notion.
If you don’t grant LLM human rights, then ClosedAI’s stance is basically that pirating works is OK because they pass them through a black box of if conditions and it leads to results that they can monetize. That’s such a solid argument, it’ll surely play well in the court of law.
Training data is not an “LLM does it”; first because “it” here is not “learning” or understanding in human sense (otherwise you would have to presume that an LLM is a human), and second because a software tool doesn’t have agency and it’s really just Microsoft using a tool based on copyrighted works to generate profit.
What I expect to happen is whoever has the most influence and power will get what they want and we'll end up raising a generation with the implicit understanding of "that's just how things are," natural order, truth, reality, and all that jazz.
The only thing that ever changes outcomes is if the contradiction status quo is incapable of being managed.
Here's an article from November 2023 that discusses this:
https://not-just-memorization.github.io/extracting-training-...
Can't you, though? I'd thought in general, it's a very important for the market to be able to do just that, otherwise everything gets gummed up in webs of exclusive contractual dependencies between established companies.
Typically providers of online databases go to some effort to stop people from sharing logins. Even from that point or view, I can imagine scraping articles and providing paraphrases of it for a fee is fishy.
All I'm saying, to some people it's obvious that the whole LLM on scraped Internet is fair use, to me it is not obvious.
Seems like the "problem" is that NYT etc gives privileged access to search engines for indexing their content, but then get upset when snippets of the indexed content is being shown to users without the users having to fight the paywall or whatever.
This article also claims that the screenshot is coming from ChatGPT when it clearly is not.
I'm not sure the problem goes away simply if the LLM in question (or any other one) gets some "no verbose regurgitation" filter.
It’s not clear to me where the line is.
If the verbatim examples that have been going around are true, that’s bad. I’d love to know more details around it — prompts used, whether that’s an old model, etc. This seems like plagiarism more than anything.
Yielding verbatim snippets of copyrighted content is a problem for OpenAI though.
So the demand to destroy those databases seems very dubious to me.
Of course later violating fair use is another issue.
As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.
This is demonstrably wrong. Many countries have both freedoms, albeit some have less strong protection than others.
They said "Most other countries" and you replied with "Many countries"
"Many" does not necessary include "most" but "most" does include "many".
No, it doesn't. If a set is of sufficiently low cardinality, “most” (in extreme cases, even “all”) of the set may not be “many”.
Most-all, in fact—Catholic Presidents of the United States have been Democrats. But it is not the case that many Catholic Presidents have been Democrats.
Most women to have served on the US Supreme Court did so only after its first 200 years. But, again, there were not many women who served on the Supreme Court only after its first 200 years.
I replace the following sentence from my previous comment:
> Most other countries don't have freedom of expression and freedom of the press, so copyright law in a different country usually lacks a unifying exception test like fair use to supplement the specific enumerated exceptions.
with the following:
Copyright law in most countries usually lacks a unifying exception test like fair use to supplement the specific enumerated exceptions in each respective country.
The rest of my previous comment remains the same.
I think you’re confusing terms of service and copyright. IANAL but what you describe sounds exactly like fair use to me, irrespective of how much you are paying NYT.
I think people severely underestimate how much they've grown accustomed to this information being freely available. It's easy to say "Well it shouldn't be available with ChatGPT," but if we actually put everything back behind a paywall and stopped people from doing things like writing blogs or newsletters that summarize the news, people here would get angry very fast.
Google has been accused for years of replacing sources with their "One Box"--the big answers at the top of the page, which are usually pulled from or corroborated by search results. They don't want you to leave the search results page (where the ads are).
Instead, I think they're paying for this:
Not only that, look at a few news articles from Tier 2 and down publications, and you'll realize that almost all of them are directly sourced from NYT and others. They'll say "so and so happened, according to The Times" (and usually link the article there)
Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying:
1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use.
2. People should familiarize themselves with the four factors of fair use determination. In particular, if a work is purely derivative of a source work and substantially negatively impacts the market for the original work, it's very likely to not be considered fair use.
A great overview is https://fairuse.stanford.edu/overview/fair-use/four-factors/
Roll back 20+ years ago on Slashdot and you'll see the exact same thing.
Copyright has been a hot button issue on the internet for decades. People end up thinking (rightly or wrongly) that they understand it without being a lawyer.
It's very possible that the example provided above is an example of fair use in some country, and that the website offering that service could be hosted there.
Quite literally, not even the lawyers or courts understand it. This is very much a "learn as you go" exercise for humanity in general at this point in time.
I only see this phenomenon speeding up. Strange times.
This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing everything they possibly can to hang on for dear life. Society needs to move on already.
Books are just strings of letters, yet copyright has still been useful to increase the volume and utility of books.
All that said, I do find the life+70y an absurdly long time.
What do you propose is the business model for artists in the absence of copyright?
We must strengthen these business models that don't depend on artificial scarcity because this number selling nonsense was over the second computers were invented. It's as dumb as asserting that you need permission to use memcpy or the mov CPU instruction.
In the US, the original copyright length was 14 years, and then 28, and eventually the lifetime of the author plus 70 years. I think the intent of the law is economically justified, but the current length is outrageous.
This is how art worked for millenia; someone commissions a chapel roof painting, someone commissions a concerto, someone commissions a statue, someone buys a chair, etc.
Artists still do this today, and there is no issue determining value beforehand. Artists list their commission prices, or their hourly costs, etc. This is a perfectly normal thing that happens everyday.
It seems to have been invented by laywers, for lawyers. Nobody else really benefits as much as they do. The whole entirety of society vs. a single profession of dubious morality.
meanwhile, most tech is moving towards subscriptions?
Art is getting paid "non-greedily". People buy a song or art piece, and then people 10 years later buy a song or art piece. That's not one person paying twice for the same song, it's two people buying the same thing.
If people still value that art for that price later, I don't see how this is a "greedy" thing. is art magically supposed to turn open source CC0 after 5 years? Tech sure doesnt work like that.
But ok so I'm a young musician. Nobody's heard of me and nobody wants to commission a concert or album. What do I do? Quit?
@Vicinity9635 It's not just lawyer greed, it creates economic fairness by preventing others from profiting from your creative work.
In the current system, artists might work for many years on a single work, or work many years perfecting their craft before anyone wants to pay for their work. Copyright gives them a way to earn money in the future that compensates them for the work they did in the past. It incentivizes creativity. Don't get me wrong, I don't think copyright is perfect, but you really ought to think more about the system you're proposing, because it's not making much sense.
(full disclosure, I’m a techie who’s gradually woken up to the idea that the tech might just be the most abused way to exploit people)
It's a bit ironic, because a lot of tech offers partial compensation in stock. Something else that really doesn't happen in games unless you work for like, the 3-4 largest studios. So they should at least understand that your compensation is not all based on labor for time worked.
that's gone out the door in the digital age. Compaies at this point have spent centuries trying to enfoce this model while witholding stuff like stock and royalties to take a part of what the company enjoys by protifting for decades off of a single (underpaid) piece of labor.
I don't exactly sympathize with a robot now trying to do the same. Pay your labor.
I don't know. Anyone funding the work is accepting a risk.
> Should they not have been paid after 1991?
They definitely should get paid for their shows and live performances. The band itself can't be copied. Artists are extremely scarce.
Their art, however, is not. Once created, the scarcity of their recordings is artificial and fundamentally time limited anyway. Even if I were extremely tolerant of copyright, I'd argue for a term of only 5-10 years maximum with absolutely no possibility of extension.
In other words, even if we accept copyright as legitimate, they sure as hell shouldn't still be getting paid for some late 80s album. They've already been adequately compensated for those creations. If they want more, they should have to keep making new stuff so that they can benefit from new copyrights which will also expire after a short time.
Creators are not supposed to be able to strike gold once and then enjoy eternal royalties. Copyright must have short time frames or it's in breach of the social contract. The reality is we're doing creators a favor by pretending that it's hard to copy their stuff so they can make some money. We do this because they assured us that eventually all of it would belong to us: works would the public domain.
The copyright industry isn't keeping up their end of the bargain. They continuously pull the rug out from under us by extending copyright to the point we'll be long dead before our culture is returned to us. It's offensive and we should all stop pretending. They need reminding that public domain is the natural and default state of all intellectual work.
> How would you predict that value before its creation (or even after)?
I'd look at the artist's past work. If there is no past work, then I don't know.
> If you're saying that only the labor has value
I'm not saying that at all. Creations are valuable. Creators are valuable. The labor of creation is valuable.
Value is assigned to stuff by humans. Obviously humans value art. The price however is given by supply and demand. The fact is that supply of intellectual works approach infinity after they are created and therefore their prices approach zero. So it makes perfect sense to assign prices to the labor of creation but zero sense to assign a price to the product of creation. Copyright is an exercise in denying reality.
> and all labor is valued equally
I definitely did not say that. All labor is different. I value some creators a lot more than others. Some creators I don't value at all.
> that sounds sort of like marxism
I must apologize if I gave that impression. I hate marxism.
Why not? The fact is that even if the album is free, there will be people paying spotify $10/month to listen to it on demand. How is it fair that Spotify can profit from it for decades to come because they offer convenience, over the artist who made the music 10 years earlier and now relinquishes their art not even a quarter into a typical career?
Copyright is abusrd now, but it's not a bad concept. I think the original copyright law of 14 + 14 worked well enough. Life expectancy increased so I'd increase it to 14 + 14 + 14 (or 10 years after the death of the original author, whatever comes first). You fund an artist for their typical career length (if they choose to extend twice) and once they are (near) retired the song is free to work off of. In the meantime you simply negotiate if you want to use their work.
Copyright means that you need to at least pay that artist you stole from in some way, which the government enforces so artists don't stop creating.
CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.
Second, you ignored the "purely derivative" bit. You have to look at to what extent the use is derivative or transformative. See https://en.wikipedia.org/wiki/Transformative_use for a bit about that. (Note, this is a legal term defined by various precedents. OpenAI can't just argue, "Turning it into an LLM is a transform, so it is transformative!") Since CliffNotes is educational and Wikipedia is nonprofit, it is relatively easy for both to qualify as transformative.
As a result your response underscores the point that was made. There are a lot of shades of grey. You really can't just seize on a couple of phrases and key points, then jump straight to the answer. You have to understand how the courts will decide, and then accept that there is an actual judgment call whose outcome depends on the judge judging.
(I'm not a lawyer, but I have had excessive exposure to them in the past.)
Is there data that supports this? I’d be interested to know what % of people who buy a Cliffs Notes have already _bought_ the original.
But, anecdotally, it's what I've seen to be the case.
Yes.
For example, Wikipedia cites many research journals that otherwise are available only by subscription.
Prior to Wikipedia, gated information centers were the norm.
The question was NOT whether it spreads information from the articles to people who wouldn't have paid for it. The question was whether it suppresses sales of the articles to people who otherwise might have paid for it.
That's a more complicated question of fact. Some people now read Wikipedia and won't buy the article. Some people encounter the reference on Wikipedia and decide to buy the article. Which happens more?
I don't have data. But publishers do. And https://scholarlykitchen.sspnet.org/2022/11/01/guest-post-wi... shows what publishers concluded.
Publishers concluded that Wikipedia references are good for sales. And so jumped on the chance to cooperate with https://wikipedialibrary.wmflabs.org/. Which is therefore able to give free access to 90% of subscription only databases to you if you can prove that you're the kind of person who is likely to add citations to Wikipedia.
Legal questions are funny like that. You have to answer the question actually asked. If you merely answer another one that sounds similar to you, your answer is generally wrong.
You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If you are actually fully paraphrasing a presentation of facts / ideas and not just altering a couple of words here and there, then there is a very strong case for non-infringement.
No, I'm not. On the contrary, I'm really looking forward to this case because I believe it will be a great test of a bunch of concepts that are totally novel in the world of copyright law as it applies to generative AI. The only things I am presenting with confidence are:
1. That anyone who declares that something is unambiguously fair use (or, contrarily, unambiguously infringing) is likely wrong. There is simply too much latitude by judges, and there have certainly been cases where a ruling went one way, only to be overturned on appeal.
2. While I certainly have an opinion on how I think this case will be decided, I'm not presenting that with unwarranted confidence. Instead, I linked that great article on the 4 factors of fair use determination because it's clear to me lots of people are saying "fair use!" on one side or the other with no understanding of the factors judges must actually consider when making a determination.
However, none of that matters in this particular thread. There are well established precedents about paraphrasing news articles and they do not support the claim you made
Remember. The NY Times does not have a record of filing frivolous lawsuits. Particularly not against companies with deep pockets. So it is almost certainly true that a lawyer who knows the law better than you thinks that this has a real chance. So you should be looking for flaws in trivial defenses that you can think up, rather than assuming that you know best.
For example take your copyright facts defense. That would be great if the NY Times was a phone book. They aren't, in addition to facts they offer analysis, editorial positions, and so on. For example I just asked ChatGPT, "In 2016, did the New York Times generally support or oppose President Trump?" I got back an answer talking about various kinds of concerns that the New York Times had, including an editorial titled, "Why Donald Trump Should Not Be President". The copy that ChatGPT needed to have to do that has a lot more than just facts in it.
Now if you paraphrased the NY Times like ChatGPT did when it answered me, you'd have a perfect fair use defense. But you aren't doing it for money, you didn't make a copy of all the NY Times, you aren't destroying the market for the NY Times, and you're legally able to own copyright in your transformed work. OpenAI is doing it for money, did copy all of the NY Times, is seriously impacting the market for NY Times articles, and ChatGPT generated text does not get a copyright.
Fair use is filled with shades of grey. Even if ChatGPT appears to do the same thing that you do, it is far less clear that OpenAI will enjoy the same level of fair use defense.
> They aren't, in addition to facts they offer analysis, editorial positions, and so on.
Those opinions and ideas are also not copyrightable. Only expressions of them are copyrightable, which is why paraphrasing facts, ideas and opinions is not a violation of copyright.
> Fair use is filled with shades of grey.
Yes, but not all those shade are equal. There is a long history of litigation showing that paraphrasing news articles is fine.
This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language.
NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations are a) an unintended failure mode of the training process, and b) from an older version of the model that has been updated to no longer regurgitate memorized text, it's not really clear how in current form GPT could possibly be considered a derivative work.
The latter is more defensible.
On the other hand, it's understandable why NYT is worried. OpenAI itself says that occupations like: Writers and Authors, Web and Digital Interface Designers, News Analysts, Reporters, and Journalists, Proofreaders and Copy Markers are "90-100% exposed" to what OpenAI is building.
I don't care that the car replaced the horse carriage because it didn't need to compensate horses nor handlers to do so. AI being the newest iteration of scraping data from artists, writers, etc. to profit millions off of is directly using the "horse handler's" work. If these LLM's threw NYT a royalty to use their articles as training material, there wouldn't be a lawsuit.
I have worked on many documentaries and any time we said “fair use” internally what we were implicitly saying is “nobody will come after us because they know that we are probably safe under fair use if this escalated.“ But again, we could never preemptively apply it. We were just anticipating potential conflict and gauging how likely it was to occur.
In the US, whether or not you make money has little to do with whether or not your use qualifies as "fair use".
That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.
> In sum, if an original work and secondary use share the same or highly similar purposes, and the secondary use is commercial, the first fair use factor is likely to weigh against fair use, absent some other justification for copying.
(P4). It’s very likely that a noncommercial secondary use would have passed under the reasoning in Warhol. I don’t understand the point you’re trying to make.But what I was arguing was that a use is not "fair use" merely because it's noncommercial in nature. I cannot make copies of movies and give them away on the street for free and successfully claim "fair use".
Weird Al has made a fantastic living copying music while only changing lyrics. He makes very heavy use of the satire plank of Fair Use.
The “commercial” test is only part of the decision criteria for Fair Use.
The "commercial" test is only a part if the criteria and not necessary, but to say it has little impact is clearly false.
"Does Al get permission to do his parodies?
Al does get permission from the original writers of the songs that he parodies. While the law supports his ability to parody without permission, he feels it’s important to maintain the relationships that he’s built with artists and writers over the years. Plus, Al wants to make sure that he gets his songwriter credit (as writer of new lyrics) as well as his rightful share of the royalties."
The fact that he could rely on fair use is separate from whether he as an artist does rely on fair use.
Not only does he get permission from the original authors, he also pays royalties to the authors despite legally not having to do so.
If that were true, I could take a band that I hate, copy all of their music note-for-note, then release an exact copy on the market and undercut them by selling their entire discography for $0.01
Fair Use requires one of several enumerated activities, including satire, education, journalism. You can’t just copy content and hope that it passes Fair Use.
Hire a lawyer if you are unsure. But at least read the Wikipedia article on the subject if you are going to talk about it.
Based upon what? You think other publishers use NYTimes articles for free without license?
Presumably, if it can remember at least a paragraph or two of each article, then surely the same would be true of any text it ingested and the model size would approach the dataset size (probably actually much larger). I don't believe this is the case at all, even searching around, I've not found any good recent examples of it regurgitating copyrighted text verbatim.
It's cool to hate AI stuff if you're a creative atm. But gotta love those generative/algorithm based PS brushes, that's still real art!
"Indeed, the opening paragraph of "A Game of Thrones" by George R.R. Martin, with the chapter titled "Bran," starts as follows:
"The morning had dawned clear and cold, with a crispness that hinted"
And then it cuts off, whether that's because OAI now have an oh shit filter or just the model had access to the first page or publicly available articles quoting the first line, I'm not sure.
I tried other chapters and random sections and it could get a sentence or two right but then hallucinated; what's more likely NYT and GRRM? That your works are being reproduced verbatim? Or that Facebook, YouTube descriptions, fan tumblrs and hell, the publicly available and multiple GoT related wikis that include a variety of passages from the books were used as training data?
"It could be fair use if conditions a, b, and c are met. Condition a means..." ;)
I think what wouldn't be covered is reproducing substantial portions of an article, especially if it's done without attribution. Tier 2 publications that fully reprint NYT or AP/Reuters articles are usually doing this via a paid News Service or Content License. See: https://nytlicensing.com/content/new-york-times-news-service...
I hope the NYT prevails here, personally. Models will (and are) currently tainted by data they should not contain and for longer term privacy concerns this needs to be addressed early and have significant consequences or we're headed towards a world where this type of technology will make our ad-targeted world seem like a much more manageable past.
You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index.
Yes, Google offers a link to the source. If OpenAI did the same, even if only 0.1% of people clicked on the links and NYTimes hardly got any revenue from it, would that make it legal in your eyes? What if they implemented a system that detected when it was outputting a verbatim copy of something and simply paraphrased it? NYTimes clearly doesn't have copyright on paraphrased versions of their articles. I think it would be pretty silly if the government forced them to do that as it wouldn't make any practical difference to anyone.
Google has a wide range of products and shakedowns. Not all of them are "perfectly" legal: Google is being challenged in court over some of their shakedowns and products practices.
Paraphrasing is also known as cloning and is often a copyright violation
In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be considered infringement if a human did it.
Genuinely - what are you talking about besides your own assumptions? you just assume everything google does is legal and therefore any one else doing anything arguably similar must also be legal? Without regard for factual details that do matter to copyright law? Such as license?? Your own description of copyright law here is very stunted - you can't paraphrase articles of the NYTimes and call it a fair use. You can report on what the NYtimes reports on... because that's what news is.
Not an assumption. This is well established. They've been doing it for twenty years!
> Without regard for factual details that do matter to copyright law? Such as license??
What license? Google doesn't in general have or need an explicit license to crawl websites and neither does OpenAI.
It's not at all well-established. How many anti-trust suits is Google facing now? Your proposition defies common sense.
>What license? Google doesn't in general have or need an explicit license to crawl websites and neither does OpenAI.
It's not the crawling the website that OpenAI did that it needs a license for... why bother conversing if you are going to be this obtuse?
Seems like the legal answer is unclear but, like Napster, such a system seems like it would lose in court.
One site clones fox news. One clones news max. And so on, cloning many news sites, sports sites, any news site. Automated, massive scale content farming. Think of the websites recommended by Taboola but, realistically, a whole lot worse.
Nobody is seriously going to ChatGPT and trying to trick it into regurgitating old NYT articles as an alternative to paying for access to NYT's archives. Meanwhile, newspapers went as far as getting the laws changed in several countries because they felt Google was competing with them too much and didn't like the fact that it was legal.
Can they? Here's reference to a legal fight where Google scraped song lyrics from a lyrics website, and presented the lyrics verbatim directly to users (bypassing the original site and the ads that allowed that site to operate)
https://www.rollingstone.com/music/music-features/genius-law...
The whole point of a search engine (as we've classically known them) is to index the web and respond to queries with a list of links that you will inspect and click through on. The whole point of an LLM chatbot tool is to eliminate those inspecting and clicking-through steps, becoming a one-stop shop for content whose substance was created by someone else. That's also the whole point of GP's hypothetical, which is why it works as an analogy.
---
There are substantially better arguments for search engines being legitimate fair use. Consider, for example, transformation. AI defenders will argue that these systems are transformative because they reshuffle elements of their input in their output, but that's clearly a much weaker form of transformation than one in which the transformed work has an entirely different nature and purpose, i.e. search engines vs. the results they return. Ultimately these technicality-based "nuh uh" arguments aren't going to save the practice of training AI on unlicensed data, because they are incompatible with the spirit of copyright law even if the novel nature of these technologies means the letter of said law can't quite nail them down yet.
If these arguments do succeed, it will be because the judicial/regulatory environment in which they were applied has been corrupted by capital.
An LLM takes an input string a corpus of text, and returns a series of text that best comes after the next input string.
To get a paragraph of output, you run the search over and over again
Both the search and LLM reshuffle the inputs to the outputs.
If I'm describing the purpose of the LLM, it's got a wide number of usages. "Making my resume look more professional" or "be a crud api" or "reformat my ask into a api call to X service" or "give me a timeline of events surrounding Y with source links"
If I'm describing the purpose of an 18 wheeler, it's got a wide number of usages. "Carry my chicken" or "carry my lettuce" or "carry my Cheetos". Or, simply, "carry my groceries".
And how did the training data contribute to the content in any meaningful way? Inspiration isn't substance.
You think all fantasy writers gotta pay Tolkien estate bc so much of fantasy draws from his tropes? Lmao no.
If training data is so unimportant, why not simply not use it and avoid the controversy? At the very least that would certainly fix the issue where the model demonstrates how "inspired" it is by NYT articles by reproducing them verbatim.
:)
Possession is not a crime when it comes to copyright. It's not like physical things (e.g. drugs or guns) at all. This is why comparing copyright violations to theft is silly.
ChatGPT can absolutely keep verbatim copies of the entire works of basically anything and not run afoul of the law. When it regurgitates a small part of an article that's covered by fair use in theory but the truth is that fair use can only be determined by a judge in a court of law when someone is sued. It cannot be determined with any sort of certainty ahead of time. It's a legal defense, nothing more.
Summarizing content has been legal forever as well (see the other posts here talking about Cliff Notes and some similar products). That's not even fair use that's just like, people's opinions, man (legally speaking).
I don't think the NYT will get what they want out of this at all.
Man thinks piracy is legal
That's not a good question.
If I look out of my window and see my neighbor go to the shop, that's fine. If I use cameras and track everybody I see on the street and put them in a database, then that's problematic and illegal in many places.
Logic does not necessarily apply when scaling is involved.
But what I'm saying is that answering the question does not allow you to deduce anything about your rights; that's what I mean by "not a good question".
If we want to establish whether scenario A is fair use or not, and we all agree that A is "worse" (regarding fair use status) than some other scenario B, then if we also agree that B is not fair use, A by definition isn't either. The opposite is not true, of course: B being fair use does not imply that A has to be as well.
I find that kind of upper/lower bound logic can be pretty useful and I think it's what the parent comment was trying to do.
On a related note, that same logic is why I think Godwin's law can be a bit misapplied now and then. Sometimes bringing up nazis/Hitler can be useful to establish some ground truth in a debate (instead of just a way to imply your opponent is actually a bad person, or, possibly, an actual nazi themselves). E.g. a conversation on the morality of violence is vastly different depending on whether you agree that violence against nazis is ok or not.
Afaik not illegal in the US. You put a camera on your own private property (window), use it to record what’s happening in a public space (the street outside), and then store that data in a database (that other people can presumably access). Unless I am missing something, this scenario is perfectly legal in the US.
Inbefore I get hit with “not every country is like that at all,” NYT is based in the US and the lawsuit is filed in the US. So how a bunch of other countries deal with similar issues shouldn’t really have as much bearing on this specific case.
So like....Wikipedia, CliffNotes, encyclopedias, etc?
None of these pay royalties to original.
> Implications: The Ninth Circuit's declaration that selectively banning potential competitors from accessing and using data that is publicly available can be considered unfair competition under California law may have large implication for antitrust law. [citation needed]
> Other countries with laws to prevent monopolistic practices or anti-trust laws may also see similar disputes and prospectively judgements hailing commercial use of publicly accessible information. While there is global precedence by virtue of large companies such as Thomson Reuters, Bloomberg or Google [or LexisNexis or Westlaw] effectively using web-scraping or crawling to aggregate information from disparate sources across the web, fundamentally the judgement by Ninth Circuit fortifies the lack of enforceability of browse-wrap agreements over conduct of trade using publicly available information.
NYT articles are largely behind a paywall for everyone. That means they are not publicly available, and a competitor who was blocked from accessing or reproducing that content without a license would not be "selectively banned"
Consider the analogy from libraries that want to do data mining.
"Unfortunately, in licenses for digital scholarly content the majority of content acquired by research libraries publishers often include terms that prohibit certain uses that would otherwise be allowable under the Copyright Act. For instance, licenses may require libraries or individual researchers to negotiate for otherwise lawful activities, such as text and data mining, and to pay exorbitant fees on top of the cost of the content itself. While new regulations allow researchers to circumvent technological protection measures to access copyrighted materials, licenses for that content may include terms that explicitly prohibit this circumvention. In many cases, these activities might actually increase the value of published material; for instance, if a data-mining project yields new knowledge about a topic covered in a journal, it may very well spark new interest in that journals content. Libraries and publishers have often assumed that license terms that restrict copyright exceptions are enforceable under state contract law. There is, however, surprisingly little case law on this point."
https://www.arl.org/wp-content/uploads/2022/07/Copyright-and...
Putting some string in a robots.txt to try to stop data collection is an amusing "solution". Should copyright owners have "Terms of Use" that limit usage for commercial "AI" purposes.
However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without payment to train LLMs is a substitutive use that is not justified by any transformative purpose."
This is a strong claim that just downloading articles into training data is what violates the copyright. That GTP outputs verbatim copies is a red herring. Hopefully the judge(s) will notice and direct focus on the interesting, high-stakes, and murky legal issues raised when we ask: What about a model can (or can't) be "transformative"?
I can't take NY Times articles, translate them into Spanish, and then sell the translations under fair use, even though clearly I've transformed the original article content.
Suppose I’m selling subscriptions to the New Jersey Times, a site which simply downloads New York Times articles and passes them through an autoencoder with some random noise. It serves the exact same purpose as the New York Times website, except I make the money. Is that fair use?
They transformed the weights.
Just like reading the article transforms yours.
As for verbatim reproduction, I'm pretty sure brains are capable of reproducing song lyrics, musical melodies, common symbols ("cool S"), and lots of other things verbatim too.
Those quotes from Dr. King's speech that you remember are copyrighted, you know?
They should be.
Times change. We're industrializing information creation and consumption (the latter is mostly here already), and we can't be stuck in the old copyright regime. It'll be useless in very short order.
All this road bump will do will give the giant megacorps time to ink deals, solidify their lead, and trounce open source. Twenty years on, the pace of content creation will be as rapid as thought itself and we'll kick ourselves for cementing their lead.
This is a transitional period between two wildly different worlds.
But nobody would do that, because ChatGPT is a really shitty way to read NYT articles (it's stale, it can't reliably reproduce them, etc.). All that is valuable about it is the way that it transforms and operates on that data in conjunction with all the other data that it has.
The real world use of ChatGPT is very transformative, even if you can trick it into behaving in ways that are not. If the courts act intelligently they should at least weigh that as part of their decision.
Suppose I start a service called “EastlawAI” by downloading the Westlaw database and hiring a team of comedians to write very funny lawyer jokes.
I take Westlaw cases and lawyer jokes and feed them to my autoencoder. I also learn a mapping from user queries to decoder inputs.
I sell an API and advertise it to startups as capable of answering any legal question in a funny way. Another company comes along with an API to make the output less funny.
Have I created a competitor to Westlaw by copying Westlaw’s works for their original expressive purpose and exposing it as an intermediary? Or have I simply trained the world’s most informative lawyer joke generator that some of my customers happen to use for legal analysis by layering other tools atop my output?
Did I need to download Westlaw cases to make my lawyer joke generator? Are the jokes a fair-use smokescreen for repackaging commercially valuable copyrighted data? Does my joke generator impact Westlaw in the market? Depends, right?
To be clear, whether the use of the original work is transformative is one key consideration within one of the four prongs of fair use. The prong "purpose and character of the use" can be fulfilled by other conditions [1]. For example, using the original work within a classroom for education purposes is not transformative, but can fulfill the same "purpose and character of the use" prong. Whether the use is for profit and to which extent are other considerations within that prong. A profit purpose doesn't automatically fail the purpose prong, and a non-profit purpose doesn't automatically pass the purpose prong.
[1] https://en.wikipedia.org/wiki/Fair_use#1._Purpose_and_charac...
This is not a RLHF problem. What I was expecting them to do is to keep a bloom filter of ngrams for known copyrighted content, such as enumerating all sets of n=7 consecutive words in an article, and validate against it. The model would only output at maximum n-1 words that look verbatim from the source.
But this will blow up in their face. Let's see:
- AI companies will start investing much more in content attribution
- The new content attribution tools will be applied on all human written articles as well, because anyone could be using GPT in secret
- Then people will start seeing a chilling effect on creativity
- We must also check NYT against all the other sources, not everything the write is original
- Paraphrasing n=7 words (and quite a few more) within a sentence can easily be fair use.
- As n gets big, the bloom filter has to also.
If/when attribution is solved for LLMs (and not fake attribution like from Bing or Perplexity) then creators can be compensated when their works are used in AI outputs. If compensation is high enough this can greatly incentivize creativity, perhaps to the point of realizing "free culture" visions from the late 90s.
I tested this 6-gram "it won't find anything matching exactly", no match. Almost anything we write has never been said exactly like that before.
This approach is probably inadequate. In my line of (NLP) research I find many things have been said exactly many, many times over.
You can try this out yourself by grouping and counting strings using the many publically available Bigquery corpora for various substring lengths and offsets, e.g. [0-16]; [0-32]; [0-64] substring lengths at different offsets.
Who pays the compensation? If it's the user, why wouldn't they just buy the authors work directly? Why go through the LLM middleman?
If it's the user, why wouldn't they just buy the DVDs directly? Why go through the Netflix middleman?
A retort to this would be that both NYT and ChatGPT are on the internet, so it's no added fuss of hopping in my car, driving to Walmart, and picking up a DVD case. My response to it would be that both the LLM and Netflix are content aggregators to the user. I can read the NYT, or I can read the NYT summary on ChatGPT and ask it for life advice with my pet hamster, or ask it how to reverse a linked list in bash.
Then there's the issue that however you credit attribution, it creates a game of enshittified content creation with the aim of being attributed as often as possible, regardless of whether the content really offered anything that wasn't out there already.
Specifically, the NYT examples all seem to be cases where they asked the AI to repeat their articles verbatim? So they ask it to violate copyright and because it's a helpful bot with a good memory, it does so.
Solution: teach the model to refuse requests to repeat articles verbatim. It's easily capable of recognizing when it's being asked to do that. And that's exactly what OpenAI have now done.
So the direct problem the NYT is complaining about - a paywall bypass - is already rectified. Now it would seem to me like the case is quite weak. They could demand OpenAI pay them damages for the time ChatGPT wasn't refusing, but wouldn't they have to prove damages actually happened? It seems unlikely many people used ChatGPT as a paywall bypass for the NYT specifically in the past year. It only knows old articles. OpenAI could be ordered to search their logs for cases where this happened, for example, and then the NYT could be ordered to show their working for the value of displaying a single old article to a non-subscriber, and from that damages could be computed. But it wouldn't be a lot.
That's presumably why the case goes further and argues that OpenAI is in violation even when it isn't repeating text verbatim. That's the only way the NYT can get any significant money out of this situation.
But this case seems much weaker to me. Beyond all the obvious human analogies, there is precedent in the case of search engines where they crawl - and the NYT let them crawl - specifically to enable the creation of a derived data structure. Search engine indexes are understood to be fair use, and they actually do repeat parts of the page verbatim in their snippets. Google once even showed cached versions of whole pages. And browser makers all allow extensions in their stores that strip ads and bypass paywalls, and the NYT hasn't sued them over that either.
This demonstrates that no, the NN actually does contain the full articles, copied into the NN. Do you think any normal person would get away with copying MS windows by e.g. zipping it together with some other OS on the same medium. Why should we let OpenAI get away with this?
> Why should we let OpenAI get away with this?
IP rights, like other private property rights, are a compromise between creators and consumers. What "should" be the case is essentially an argument about what balance creates the best overall outcomes. LLMs, for now, require large amounts of text to train, so the question is one of whether we want LLMs to exist or not. That's really a question for Congress and not the courts, but it'll be decided in the courts first.
The entity which owns ChatGPT is apparently maintaining a copy of the entirety of the New York Times archive within the ChatGPT knowledge base. That they extract some fair use snippets (they would claim) from it would still be fruit of a poisoned tree, no?
(disclaimer: I'm pro AI, anti copyright, especially anti elitist NY Times; but pro rule of law)
Your creative work does deserve at least some period of exclusive rights for you. Definitely not so much that your grandchildren get to quibble about it well into retirement. But also whatever the number 3 or 4 most valuable company in the world doesn’t get to scrape your content daily to repackage and sell as intelligent systems.
Here's a thing though: for 99%+ of that content, being turned into feedstock for ML model training is about the only valuable thing that came of its existence.
If it were not for world-ending danger of too smart an AI being developed too quickly, I'd vote for exempting ML training from copyright altogether, today - it's hard to overstate just how much more useful any copyrighted content is for society as LLM training data, than as whatever it was created for originally.
Note that I don't see any major problem if only articles that were, say, more than 5 or 10 years old were being used. I don't think the current length of copyright makes any sense. But there is a big difference from last year's archive vs today's news.
This can be further coupled with search - use GPT to look at multiple sources at once, and report. It's what humans do as well, we read the same news in different sources to get a more balanced take. Maybe they have contradictions, maybe they have inaccuracies, biases. We could keep that analysis for training models. This would also improve the training set.
LLMs are arguably compressed data archives with weird algorithms. The fact that they will regularly regurgitate verbatim quotes of training data is evidence of this, as are the guardrails that try to prevent this.
The second piece of evidence is this paper explained here https://www.hendrik-erz.de/post/why-gzip-just-beat-a-large-l... where instead of an LLM researchers used gzip compressed data as a model and it even beat trained LLMs.
AI is a bit of a black box, but that doesn’t protect the operators of black boxes from rights violation suits. You can’t make a database of scraped copyrighted data and patented that querying that data is fair use.
There needs to be law made here and the law just isn’t going to be “everybody can copy everything for free as long as it’s for model training”.
Licensing will have to be worked out, actual laws and not just case law needs to be written. I have a lot of sympathy for lots of leeway for the open source researchers and hackers doing things… but not so much for Microsoft and Microsoft sponsored openai.
Further, the evidence presented by NYT in the lawsuit could be hard to reproduce. I tried multiple prompts on multiple versions of GPT-4 APIs but still could not get GPT-4 to reproduce NYT articles exactly. NYT might as well tried to let GPT-4 reproduce 100,000 articles and only found a few cases where GPT-4 actually recited the whole article. In that case OpenAI might as well be arguing that this is only a rare bug and avoid losing the lawsuit in a massive way.
OpenAI has created a $100bn company on this transfer. The Times may have an interest in a material fraction of that wealth.
The Times almost certainly wants its own LLM. I could see them striking a consortium agreement with other newspapers more easily than OpenAI.
OpenAI alone has a market cap that'd allow it to buy about as large a proportion of publishers of newspapers and books as they'd be allowed before competition watchdogs will start refusing consent.
Put another way:
If I was a VC with deep pockets investing in AI at this point, I'd hedge by starting to buy strategic stakes in media companies.
I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder.
Kinda funny if it works though. We'd first train them to copy their training data verbatim, then train them not to.
That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's just that they're given so much data that we don't expect this to be possible for most of the training data given the parameter count.
One thing you might do is use a full-text search database of the entire training data. If part of ChatGPT response is directly copied, give it the assignment of "please paraphrase this" and substitute the paraphrase into the response. This might slow ChatGPT down a lot - but it might not, I think an LLM is actually more computationally expensive than a full-text search by a lot.
Production open access LLMs do probably need a front-end filter with a fine tuned RAG model that identifies and prevents spitting out copyrighted material. I fully support this.
But we shouldn't be preventing the development of a technology that in 99.99% of usecases isn't doing that and can used for everything from diagnosing medical issues to letting coma patients communicate with an EEG to improving self-driving car algorithms because some random content producer's works were a drop in the ocean of content used to learn relationships between words and concepts.
The edge cases where a model is rarely capable of reproducing training data don't reflect infringement of training but of use. If a writer learns to write well from a source is that infringement? Or is it when they then write exactly what was in the source that it becomes infringement?
Additionally, now that we can use LLMs to read brain scans and have been moving towards biological computing, should we start to consider copying of material to the hippocampus a violation of the DMCA?
Isn't that in tension with the basic idea of an LLM of predicting the next token? How do you achieve that while never getting close enough to plagiarism?
I feel like the NYTimes is asking for deletion as a negotiation tactic to force OpenAI to give them enough money to pay for their journalism (I am not sure who would subscribe to NYTimes if you can get as much through OpenAI, but I am open to registering extra to pay for their work).
It's the other way around. There is no infringement if the model output is not substantially similar to a work in the training set [1]:
> To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright, the defendant actually copied the work, and the level of copying amounts to misappropriation.
The questions are, which parties should bear liability when the model creates infringing outputs, and how should that liability be split among the parties? Given that getting an infringing output likely requires the prompt to reference an existing work (which is what's happening in the article), an author of a work, an element in an existing work, or a characteristic/style strongly associated with certain works/authors, I believe that the user who makes the prompt should bear most of the liability should the user choose to publish an infringing output in a way that doesn't fall under fair use. (AI companies should not be publishing model outputs by default.)
[1] https://en.wikipedia.org/wiki/Substantial_similarity#Substan...
Its true that OpenAI will defend the wholesale copying into the training set by arguing that the transformative purpose of the next use reaches back and renders that copying fair use, but while that's clearly the dominant position of the AI industry, and it definitely seems compatible with the Cobstitutional purpose of Fair Use (while currently statutory, the statutory provision is codification of Constitutional case law), it is a novel fair use argument.
NY Times is suing because of both the model outputs and the existence of the training set. But infringement in the training set doesn't necessarily mean that the model infringes. Why? Because of the substantial similarity requirement. But first, I'll address the training set.
For articles that a person obtains through legal methods (like buying subscriptions) but doesn't then republish, storing copies of those articles is analogous to recording a legally accessed television show (time-shifting), which generally is fair use. Currently, no court has ruled that "analogous to time-shifting" is good enough for the time-shifting precedent to apply, but I think the difference is not significant. The same applies to companies. Companies are not literally people, but there isn't a reason for the time-shifting precedent to not apply to companies.
What about the articles that OpenAI obtained through illegal methods? Then the very act of obtaining those articles would be illegal. The training set contains those copies, so NY Times can sue to make OpenAI delete those copies and pay damages. But it's not trivially obvious that a GPT model is a copy of any works or contains copied expression of the any works in the training set; the weights that make up the model represent millions of works, it's not trivially obvious that the model contains something substantially similar to the expression in a work in the training set. Therefore, it's not trivially obvious that infringement with respect to the training set amounts to infringement with respect to the model made from the training set. If OpenAI obtained NY Times articles through illegal means, then making OpenAI delete the training set would be reasonable, but the model is a separate matter.
As long as the model doesn't contain copied expression and the weights can't be reversed into something substantially similar to expression in the existing works, then what matters is the output of the model.
If a user gives a prompt which contains no reference to an existing NY Times author, work, or a strongly associated characteristic/style, then do OpenAI's models produce outputs substantially similar to expression in the existing works? If not, then OpenAI shouldn't be liable for infringing works, because the infringing works result from the user's prompts. If my premise is false, then my conclusion falls apart. But if my premise is true, then at most I would admit that OpenAI has a limited burden to prevent users from giving those prompts.
The thing about you claim, "Just learn to recognize and punish plagiarism via RLHF" is that we've had an endless series of prompt exploits as well as unprompted leakage and these demonstrate that an LLM just doesn't have fixed border between its training data and its output. This will it basically impossible for OpenAI to say "we can logically guarantee ChatGPT won't serve your data freely to anyone".
1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to regurgitate.
2. Research/hosting/progress will proceed. The US cannot stop this, only choose to be left behind. The world will move on, with China gleefully watching as their biggest rival commits intellectual suicide all to appease rent seeking media companies.
3. Models can share weights, merge together, cooperate, ablate, evolve over many generations (releases), etc. Copyright law is woefully ill equipped to handle chasing down violators in this AI lineage soup, annealed with data of dubious/unknown provenance.
I could go on, but the point is that, for better or worse, we live in a new intellectual era. The NYT et al are coming along for the ride, whether they like it or not.
Analyzing the factors involved for a "fair use" consideration:
Purpose and Character of the Use: While the argument for transformation might hold in the future as you point out, the current dispute revolves around verbatim use. So clearly not transformative. Also commercial use is more difficult to be ruled fair use.
Nature of the Copyrighted Work: Using works that are more factual may be more likely to be considered fair use, but I would argue that NYT articles are as creative as factual.
Amount and Substantiality of the Portion Used: In this case, the entirety of the articles was used, leaving no room for a claim of using an insignificant portion.
Effect on the Market Value: NYT isn't getting any money from this, and it's clearly not helping their market value if people are checking on ChatGPT instead of reading a NYT article.
IANAL, but in my opinion NYT is well within its rights to pursue legal action. Progress is inevitable, but as humans, we must actively shape and guide it. Otherwise it cannot be called progress. In this context, legal action serves as a necessary means for individuals and organizations to assert their rights and influence its course.
In the case of the famous screenshot, the AI just relayed the information it found on the web, it's not included in its training data.
So you're just wrong.
Rather, verbatim reproduction is the proof that copyrighted materials was used. Then the court has to evaluate whether it was fair use. Without verbatim reproduction, the court might just say that there is not enough proof that the Times's work was important for the training, and dismiss the lawsuit right away.
Instead, the jury or court now will almost certainly have to evaluate OpenAI's operation against the four factors.
In fact, I agree with the parent that ingesting text and creating a representation that can critique historical facts using material that came from the Times is transformative. An LLM is not just a set of compressed texts, people have shown for example that some neurons fire when you are talking of specific historical periods or locations on Earth.
However, I don't think that the trasformative character is enough to override the other factors, and therefore in the end it won't/shouldn't be considered fair use IMHO.
What if a human manually searches all those articles and transcribes / summarizes them to me in the way ChatGPT did?
People are not using ChatGPT as a replacement for current news, and because of hallucinations, no one should be using it for past news either. I wouldn't remotely call ChatGPT a competitor of NYT traffic, like I would Reuters or other news outlets.
Because if it is not good enough, then it is not a market substitute.
The laws cares if it is a market substitute and if there are damages. If it sucks, then there aren't damages, which matters for the 4th factor of fair use.
Rent seeking? Media companies that actually create content are rent seeking? Versus the garbage hallucinations AI creates?
I know because they tried to make a deal with my company, we passed because social media data is infinitely more valuable.
Maybe their data isn't as valuable to eg. advertisers than the data their audience actually shouted into the internet themselves (guess what), but the thing they've been actually selling for a long time now, journalism, can't be dying that fast considering we're both on this website that in big parts consists of discussing journalism.
> ”Rent seeking” is one of the most important insights in the last fifty years of economics and, unfortunately, one of the most inappropriately labeled. Gordon Tullock originated the idea in 1967, and Anne Krueger introduced the label in 1974. The idea is simple but powerful. People are said to seek rents when they try to obtain benefits for themselves through the political arena. They typically do so by getting a subsidy for a good they produce or for being in a particular class of people, by getting a tariff on a good they produce, or by getting a special regulation that hampers their competitors. Elderly people, for example, often seek higher Social Security payments; steel producers often seek restrictions on imports of steel; and licensed electricians and doctors often lobby to keep regulations in place that restrict competition from unlicensed electricians or doctors.
https://www.econlib.org/library/Enc/RentSeeking.html
This is linked in the wikipedia article, which is even more confused:
Indeed, it is quite different, because those things are scarce physical things in the real world. Intellectual property is a scam, and killing it once and for all will be one of the best things to come out of the current AI hype cycle. Nobody will "own" ideas, pieces of information, or strings of bytes.
Is that just by increasing the temperature, tweaking the prompt, etc.? If you can operate on the raw weights and recreate the original text, copyright infringement still applies.
Sorry, is this the same China that has already introduced their own sweeping regulations on AI? Which in at least one instance forced a Chinese startup to shut down their newly launched chatbot because it said things that didn't align with the party's official stance on the war in Ukraine?
https://finance.yahoo.com/news/beijing-tries-regulate-china-...
https://nitter.unixfox.eu/CDT/status/1625936306814717952?337...
I don't disagree that research/hosting/progress will continue, but I'm not so sure that it's China who stands to benefit from the US adding some guardrails to this rollercoaster.
Your second point reminds me a bit of 'War with the Newts' where humanity arms a race of sentient salamanders until they overthrow humanity. How could we not arm our newts if Germany might be arming theirs?
I also think basically everything else you wrote is wrong.
You don't have to agree with it. You don't have to like it. But if you accept it and live by it, it's much harder to get burned.
There are plenty of large companies in other sectors that acknowledge there are limited legal remedies for them if someone copies some aspect of their business or name.
https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
From page 30 and onwards has some fairly clear examples on how ChatGPT has an (internal) copy of copyrighted material which it will recite verbatim.
Essentially if you copy a lot of copyrighted material into a blob and then apply some sort of destructive compression to it. How destructive would that compression have to be for the copyright no longer to hold? My guess it would have to be a lot.
As I see it the closeness of OpenAI may be what saves it. OpenAI could filter and block copyrighted material from the LLM from leaving the web interface using some straight forward matching mechanism against the copyrighted part of the data set ChatGPT has been trained on. Whereas open source projects trained on the same data set would be left with the much harder task of removing the copyrighted material from the LLM itself.
I imagine the goal is closer to "enough that no one notices we stole it", either in a way that it's not easily discoverable or even when directly analyzed there's enough plausible deniability to scrape by.
It makes it difficult for me to ascertain whether it is repeating from it's training data, or they committed the same mistake as the OP article of using Copilot, which ends up googling(binging?) the article first, before replying.
I hope a court establishes some rules of engagement here, even if it’s not this case.
FWIW, I can’t replicate on either GPT 3.5 or 4, but it may be that OpenAI has added new measures to prevent this.
That said, the meme image of Willy Wonka comes out of stable diffusion 1.5 almost perfectly with surprising frequency. Then again, this is probably because it appeared hundreds or thousands of times in the training set in all sorts of contexts because it's such a popular meme. There is a tension between its status as an integral part of language and its nature as a copyrighted screen grab.
However, I had good luck reproducing poems on GPT 3.5, both copyrighted and not copyrighted, because the choice of words is a lot more "specific" so to speak, and therefore higher temperature isn't enough to prevent complete reproduction of the originals. See https://chat.openai.com/share/f6dbfb78-7c55-4d89-a92e-f4da23... (Italian; the second example is entirely hallucinated even though a poem with that title exists, while the first and third are recalled perfectly).
I’m more surprised that it can repeat 100 articles; if that behaviour is consistent in larger sample sizes and beyond just NYT dataset (which might be repeated on the web more than other sources, causing overfitting), that would be impressive.
You could imagine at some point a large enough GPT5 or 6 or 7 will be able to memorize verbatim every corner of the web.
It's more like, is the new work a distinct expression, e.g. satire or commentary, based on the original.
You can reproduce the original verbatim and still be transformative by adding an element of critique.
Example: https://www.dmca.com/articles/akilah-obviously-vs-sargon-of-...
in japan, where they said anything goes for ai
so its best to not to lose a competitive edge with things that people openly publish on the internet, if you put it out there for everyone to see then expect other people to use it
its about a precedent. If you don't keep up with international competition, you lose.
This would be a huge blow to open-source and research developers and I'd even argue it could help openAI to get a bit of a moat ala regulatory capture.
Google won that suit under fair-use as a massive searchable database was found to be transformative as well as the non-commercial nature.
So; if your web scraping companies goal is to allow people to bypass a paywall I suspect you'll have trouble in the future. If your web scraping company instead say allows people to do market analysis on how many people need a piano tuner in NYC and it doesn't do that by copying a NYT article doing original research I think you'll be fine.
I think all this is doing is making us realize that we have built a massive economic system on a fundamentally flawed idea of ownership over ideas, and the only two solutions will be to tear up the rule book, which will be extremely painful, or double down, which will be fatal.
But they are not. It's much simpler, proprietary writing is now integrated into the source code of OpenAI, it would be as if I would copy parts of other propriety code and copy paste it into my own codebase. Claiming copy paste is a natural evolving process of millions of years of evolution.
The fact that LLM's are so complicated and we don't know where it is, doesn't make it less so.
It's not copy-pasted; it's compressed in a lossy manner. Even GPT4 has nowhere near enough memory to store the entirety of its training data in a non-lossy compression format. Just likes how humans compress the information we read.
Humans don’t have the scale machines have and moreover humans aren’t sevices, that argument doesn’t fly.
I really think NYTs data isn’t that important and nor crucial, LLMs could’ve just elided it. However, it’s more about training on copyrighted data in general which is kind of crucial for OpenAi, they trained their LLMs indiscriminately on copyrighted content without any plan to share any profits.
Let alone that it's a centralised model that's being distributed for a fee.
Software programs are not humans, and need to be treated differently. Anthropomorphization is one of the slipperiest paths to argue anything.
If only small patches of the original image can be reproduced then it becomes much more murky.
For example, the Wikipedia article
https://en.wikipedia.org/wiki/List_of_people_claimed_to_poss...
contains several examples of people who were able to look at pages and recite them back. That is actually a much stronger ability than GPT since GPT has presumably looked at them 100 times.
Laws are created by people (not by computers reasoning that all analogies must be true). And fairness is an important part of that process.
Developers thinking LLMs are akin to humans arent the brightest crop, and are usually a topic of ridicule.
The source code of the LLM is likely a few hundred lines of text describing the shape of the neural networks involved in the model.
None of the NYTimes content will be in the source code. NYTimes doesn't publish Python source code, it publishes human language news.
LLMs are conceptually simple, mostly matrix multiplications and some non-linear operations connecting each layer, in some loops based on attention, etc. It's the staggering amount of training data and compute that makes them complex.
NYT won't mind if you use their content to train LLMs - as long as they get a commission. Reddit will shut down their free API and make you pay to get training content. Discord is going to be selling content for AI training too - if they haven't already done so. Twitter is doing it.
They didn't care before because LLMs were just experiments. Now we're talking trillions of dollars of value.
Can you make the argument this was their fault for not having forward vision/being asleep at the wheel and "accidentally, in hindsight" letting OpenAI/others have free, open, unlimited access to their content?
They are not giving it out "for free", in fact they're being paid by their employer to write these articles. Moreover, the writers themselves stand noth' to gain from their past writings financially as they don't belong to the ownership structure of the business.
This is a dumb argument. We're not just talking about ancient articles. We're talking about new content, including content that is yet to be written.
ChatGPT isn't competing with NYT on a core competency. No one uses LLMs for original news reporting. They're obviously incapable of doing that, by virtue of not being there on the scene or able to independently research a topic, maintain relationships with sources, etc. What ChatGPT can do is quote/reproduce some parts of past articles, and reason from them. Or at least produce new text that's somewhat related to the old text.
The threat to NYT is this: ChatGPT is much better bullshitter than they are, so it reduces NYT to its core competency: providing original information. Which is all it should be doing in the first place. But instead, NYT wants to not only keep the bullshitting part of its revenue, but also take a cut or destroy the much greater and much more useful part of where this all feeds a general-purpose language model.
This is a badly-formulated conjecture, or worse, ultimately selective reading of "social credit" which only purpose is serving your argument; it has nothing to do with economics. I'm sorry, but I'm not convinced.
OpenSource developers did that ;)
Earnestly I found ";)" deeply troublesome.
You're fighting a scarecrow that doesn't exist...
Can you say the same for user created content on Reddit, Twitter, or Facebook? A user agreement that nobody reads doesn't have anything like the same legal basis as a signed contract. Not to mention that a large percentage of users are not adults.
Why should the law treat a LLM in a body reading NYT on a tablet differently than a LLM browsing the content from a website online and reading that?
While harder to do as a human, if memorised a copyrighted book and then did a live reading on TV, or produced replicas from memory and sold them (the most comparable example), I’d be sued.
Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you can’t do it verbatim.
This is not the most comparable example, because it's not what ChatGPT is doing. The most comparable example is if you were hired as a contractor and the employer asked you to write verbatim some copyright content you'd memorised. If the employer then published it, they'd be the one liable, not you.
>Humans produce derivative work all the time, and it’s fine for LLM’s to do that, but you can’t do it verbatim.
Nobody's suggesting preventing humans from consuming any copyrighted content just because in future they might recite some of it verbatim, but that's what NYT want for LLMs.
No, you'd both be liable. You are not allowed to create copies of a copyrighted work, even from memory, for any commercial purpose. Making it public or not is irrelevant.
This is more obvious with spftware: if I copy a version of AutoCAD that my previous employer bought and sell it to another company, or even just use it for my current employer without showing it to anyone else, I am violating the copyright on that software, and I am liable. Even though obviously no "publishing" happened.
Similarly, if you hire a decorator to paint Mickey Mouse on the inside walls of your private kindergarten, the decorator is violating Disney's copyright just as much as you are, even if neither of you has made that public.
That's the point at which infringement occurs in your example. It's not the memorizing that's the infringement, it's the reproduction from your memory.
We shouldn't be regulating your hippocampus encoding the book, but your reproducing the book from that encoding.
Similarly, we shouldn't be regulating the encoding of material into the NN, but the NN spitting back out the material.
Are they all owned by one mega-corporation, which is going to do as capitalism does, and use them to squeeze money out of all of us? Then I'm happy to ban them.
The opportunity cost of holding this technology back is going to literally be millions of people's lives given current trends in its emerging applications.
Police usage, not training.
You'd get the same problem with someone with a photographic memory who a group of people would turn to recite them the news instead of buying the newspaper.
As of now public performance of copyrighted material is infringement.
I fully agree with the perspective that infringement in usage needs to be limited even if I strongly disagree that training is infringement.
This tired 'fair use' excuses from AI bros whilst the GPT has reproduced the article text verbatim, word for word and it being monetized without the permission from the copyright holder and source (NYT) is an obvious copyright violation 101. Full stop.
Again, just like Getty v. Stability, this copyright lawsuit will end in a licensing deal. Apple played it smart with their choice with licensing deals to train their GPT [0]. But this time, OpenAI knew they could get a license to train on NYT articles but chose not to.
[0] https://9to5mac.com/2023/12/22/apple-wants-to-train-its-ai-w...
the purpose and character of the use
the nature of the copyrighted work
the amount and substantiality of the portion taken
the effect of the use upon the potential market.
Literally every single one of these factors has very complicated precedent and each one is an open question when it comes to AI. Since fair use is a balancing test this could go any way.Stability took the easy way out because they didn't have billions of dollars to play around with and Microsoft to back them. Let's see what OpenAI does but calling everyone who disagrees with your naive interpretation of fair use "AI bros" is doing everyone a disservice.
What (or whom) do you consider to be an "AI bro?"
This sort of ad hominem generalization usually accompanies a weak argument.
And even if they did it will be fine because those sources allow for it.
The point is that OpenAI never asked NYT for permission to use their data.
Fair use has nothing to do with reproducibility. LLMs are more clearly fair use than a search engine cache and those court cases are long settled. There's no world in which OpenAI doesn't win this entire thing.
Why do you think the architecture is important? If I have a computer program and it outputs the an entire copyrighted poem then the answer to "is this copyright violation" SHOULD NOT depends on the architecture of the program.
It gets harder to stand behind a blanket claim that LLMs or any AI we’ve got falls under fair use when they keep repeatedly reproducing complete and identifiable individual works and clearly violating copyright laws in specific instances. The models might be remixing and/or transformative most of the time, but we have proof that they don’t do that every time nor all the time… yet. Maybe the lawsuits will be the impetus we need to fix the AIs so they don’t reproduce specific works, and thus make the fair use claim solid and actually defensible?
https://www.youtube.com/watch?v=eUHBPuHS-7s (the original is flash and has thus been consigned to the memory hole, so we are left with this poor-quality conversion)
36": 'however, the press as you know it has ceased to exist'
40": '20th-century news organizations are an afterthought; a lonely remnant of a not-too-distant past'
2'11": 'also in 2002, google launches google news, a news portal. news organizations cry foul. google news is edited entirely by computers'
5'13": 'the news wars of 2010 are notable for the fact that no actual news organizations take part. googlezon finally checkmates microsoft with a feature the software giant cannot match: using a new algorithm, googlezon's computers construct new stories, dynamically stripping sentences and facts from all content sources, and recombining them. the computer writes a new story for every user'
5'55": 'in 2011 the slumbering fourth estate awakes to make its first and final stand. the new york times company sues googlezon, claiming that the company's fact-stripping robots are a violation of copyright law. the case goes all the way to the supreme court'
they didn't get the details exactly right, but overall the accuracy is astounding
however, that may be a hyperstition artifact in this timeline
https://en.wikipedia.org/wiki/EPIC_2014 (i thought epic 2014 might be the only flash video to hae a wikipedia article about it, but then i looked and found five others)
This is interesting. The NYT is specifically saying that the way you use an LLM impacts what you can legally use for training the LLM. They're firing shots at the big guys trying to sell access to an LLM, but not at the little guy self-hosting for fun or academics doing research.
https://www.npr.org/2023/05/18/1176881182/supreme-court-side...
Down at the bottom of the linked PDF are some more interesting allegations:
Count 5 - MS/OpenAI removed NYT copyright notices in violation of the DMCA.
Count 7 - By attributing hallucinated garbage to NYT, MS/OpenAI is diluting NYT trademarks in violation of US Trademark law.
I admit: I laughed. This will be an entertaining lawsuit to follow.
The courts are going to rule that LLM training is a transformative use case that is protected as fair use under copyright law. They may rule that if an LLM-powered service is explicitly designed to enable copyright violation that is illegal, but there is no way any court is going to look at these examples and see it as anything other than the NYT fishing to try and generate a violation by using the LLM in a way that is very different than the service is intended to be used and which -- even if abused -- doesn't hurt the business model under which the text has been produced.
The most likely outcome is that LLM providers will add some sort of filter on output to prevent machines from regurgitating source documents. But this isn't a court case the NYT can win without gutting fair use protections, and that would be a terrible thing.
$750 [1] * 66 million records [the lawsuit] is basically 50 billion.
[1]: https://www.ce9.uscourts.gov/jury-instructions/node/706
https://www.newyorker.com/books/page-turner/rethinking-the-l...
The US constitution says, The Congress shall have Power
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
So the Congress's power to make copyright and patent laws is predicated on promotion of science and useful arts (I believe this actually means technology). In a sense, the OpenAI being the forefront of our AI technology advancement is crucial to the equation. To hinder the progress by copyright is, in my mind, unconstitutional.
I think we all agree that no one is entitled to “progress of science” at any cost - as a straw man, killing hundreds of newborn babies for scientific research is not great - so we use ethics and the legal system to find the line of what’s acceptable.
I don’t know exactly what NYT is asking for here, but the two options aren’t unconsented training vs nothing at all. NYT could license, for a fee, its content to OpenAI. It’s pretty common for scientists to have to pay for materials!
- Read 20 different news websites and their story on the same event/topic
- Wait an hour, grab a cup of coffee
- Sit down to write my article, never from this point I open any of the 20 news websites, I write the story from my head
- I don't consult any other source, just write from my memory, and my memory is, let's say, not the best one, so I will never write more than 10 words exactly as they appear on any of the 20 websites.
- I will probably also write something that is not correct or add something new because, as I said, my memory is not the best.
Is that fair use? Am I infringing on copyright?
Culturally we’re taught that there is a moral component to copyright and patent law - that stealing is stealing. But the idea that words or thoughts or images can be owned (and that the might if the state can be brought to bear to enforce it) would seem utterly ludicrous to someone from an earlier era. Copyright and patent laws exist for practical, pragmatic reasons - and seemingly they have served us well, but it’s not unreasonable to re-examine them from first principals.
This rings similar.
Is there any research into how people from earlier eras thought about it? And should all laws that seemed ludicrous to someone from an earlier era be discarded? If not, how exactly do we determine the relevance of what someone from an earlier era would think about our laws?
Then LLMs would be distributed only via torrents, like most copyright infringing media.
Why?
The LLM genie is out of the bottle: an unfavorable court ruling in a single country isn't going to stuff it back in.
On the other hand, if LLM are used to "launder" copyright content and, accepting the premises of copyright law, this has the effect of reducing incentives to do creative work, that has obvious negative implications for economic productivity.
Assuming this is in good faith: the ability to write code, documentation, and tests is absolutely a productivity enhancer to an existing programmer. The code snippets from a dedicated tool like copilot are of very usable quality if you're using a popular language like Python or JS.
Loading data to which you have no rights over into your software is legally perilous, yes.
It's as easy as simply asking for and receiving permission from the data's rightsholders (which might require exchange of coin) to make it not legally perilous.
I suspect it wouldn't be too hard to convince the EU though, the EU has an history of giving up rights and markets to big copyright holders even if that hurts the local companies.
LLM's will become more expensive and less attractive as money printers, this will screw with the business models of the direct provision folks like OpenAI, MS and Google, MS and Google will only shed tears for money spent while OpenAI will just not have as good an income stream until they think of something new.
I'm sure that's what they want, but I'm not sure that's what the outcome will be. What if they want to charge a prohibitive amount of money for their content?
I think Spotify vs Napster is a good example, content creators in news (the Journalists) are already in a hard place (vs. successful rock stars preinternet) I think that the news providers are rather like the music lables.
So you will use a Chinese AI that spies on you, or you will use some shady service from a shady country (that will play cat and mouse like torrent sites).. or most likely you will run your own model when you are computer literate and no model if you are not.
Actually most models are so lobotomized allready that probably better to run your own, as long as you have a good enough computer.
No need to emigrate!
The thing about lawsuits is that you make dozens of claims, and the court can rule in favor of some of them, and against others. The question of "is LLM training fair use?" hasn't made it to a high court yet. The court could very easily rule against everything else in the suit.
It is the specific use of article photocopies to circumvent the normal sale of newspapers that becomes illegal.. and even that is questionable. If I read the newspaper left out in a waiting room and it keeps me from buying that days paper, this is not a criminal act.
It’s a four part test. Let’s examine it thusly:
1. Transformative. Is it? It spits out informative text and opinion. The only “transformation” is that its generative text. IMO that’s a fail.
2. Nature of the work - it’s being used commercially. Given it’s being trained partially on editorial, that’s creative enough that I think any judge would find it problematic. Fail on this criteria.
3. Amount. It looks like they trained the model on all of the NYT articles. Oops, definite fail.
4. Effect on the market. Almost certainly negative for the NYT.
IMO, OpenAI cannot successfully claim fair use.
It is not the NYT making the claim of Fair Use, it is OpenAI.
My guess is that the court will likely find in the Times favor, because the legal system won't be able to understand how training works and because people are "scared" of AI. To me, reading a book, putting it in some storage system, and then recalling it to form future thoughts is fair use. It's what we all do all the time, and I think that's exactly what training is. I might say something like "I, for one, welcome our new LLM overlords". Am I infringing the copyright of The Simpsons? No.
I am guessing some technicality like a terms-of-use violation of the website (avoidable if you go to the library and type in back issues of the Times), or storing the text between training sessions is what will do OpenAI in here. The legal system has never been particularly comfortable with how computers work; for example, the only reason EULAs work is because you "copy" software when your OS reads the program off of disk into memory (and from memory into cache, and from cache into registers). That would be copyright infringement according to courts, so you have to agree to a license to get that permission.
I think the precedent on copyright law is way off base, granting too much power to authors and too little to user. But because it's so favorable towards "rightsholders", I expect the Times to prevail here.
Anyway, like I said, I don't think OpenAI will win this. Someone will produce one verbatim article and the court will make OpenAI pay a bunch of money as though every article could be reproduced verbatim, and AI in the US will be set back that many billion dollars. It probably doesn't matter in the long run; it preserves the status quo for as long as the judge is judging and the newspaper exec is newspaper exec-ing. That's all they need. The next generation will have to figure out how to deal with AI-induced job loss... and climate change. Have fun, next generation!
"Its what we do all the time" is a major assumption
So, if you setup a service like ChatGPT but powered by humans responding real time to queries, and these humans would occasionally reproduce large chunks of NYT articles, they and the service itself would be liable for copyright infringement. Even if they were all reproducing these from memory.
Now, this is somewhat different from the discussion of whether training the model on the copyrighted data, even if it had effective protections from returning copies of it, constitutes copyright infringement in itself. I believe this is a somewhat novel legal question and I can think of no direct corollaries.
I certainly don't think we can just handwave and say "at some level, when a human reads a copyrighted work, they are doing the same thing", because we really don't know if that is true. Artifical neural networks certainly have no direct similarity with the neural networks in the brain as far as we can tell. And, even if they did, there is no reason to give a machine the same rights that a human has - certainly not until that machine can prove sentience.
To put it another way, let's say I turn the dial all the way the other way, I train the worlds crappest LLM on NYT material, it massively massively overfits and all it will ever return is verbatim snippets of the NYT. Is that copyright infringement?
The core part of the argument here is actually just that OpenAI doesn't want to adhere to what the current standard is for using copyrighted material, if you want to use it and create something new with it you need to license the material. Since OpenAI's LLM isn't actually like a human it needs to license such a vast dataset that it would be uneconomical to run the business without stealing all the content.
NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 - Dec 2023 (80 comments)
The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 - Dec 2023 (837 comments)
The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/item?id=38781863 - Dec 2023 (11 comments)
New York Times Sues Microsoft and OpenAI, Alleging Copyright Infringement - https://news.ycombinator.com/item?id=38781718 - Dec 2023 (1 comment)
New York Times sues Microsoft and OpenAI over copyright infringement - https://news.ycombinator.com/item?id=38781908 - Dec 2023 (2 comments)
New York Times sues OpenAI, Microsoft for using articles to train AI - https://news.ycombinator.com/item?id=38782510 - Dec 2023 (1 comment)
New York Times sues OpenAI, Microsoft for allegedly infringing copyrighted work - https://news.ycombinator.com/item?id=38783699 - Dec 2023 (1 comment)
New York Times sues OpenAI, Microsoft over use of its stories to train chatbots - https://news.ycombinator.com/item?id=38784914 - Dec 2023 (1 comment)
NY Times sues OpenAI, Microsoft for infringing copyrighted works - https://news.ycombinator.com/item?id=38786330 - Dec 2023 (1 comment)
NYTimes sues OpenAI, Microsoft, for copyright infringement - https://news.ycombinator.com/item?id=38790845 - Dec 2023 (1 comment)
If the AI can recall the text verbatim then it's not at all the same. When we read we are not able to reproduce the book from our memory. Even if a human could memorise an entire book it's not at all practical to reproduce the book from that. The current AIs are not learning "ideas", they are learning orders of words.
However I am inclined to agree with them for the simple fact that putting a file into a device and letting that device reproduce parts of the file should be allowed. I mean we're already at the point where this simple right is under pressure from DRM, but people should be allowed to do whatever they want with the files they own.
Whether you can publish this output and share it with the world is a whole different issue.
[0] https://www.nytimes.com/2023/12/22/technology/apple-ai-news-...
[1] https://www.theverge.com/2023/12/22/24012730/apple-ai-models...
Unless they engage in massive IP and DNS banning, geolocation based, that forced upon all internet users and "external" users.
If a NYT article says "Henry Kissenger was known to eat ice cream on a hot day" and our game outputs the same, it is purely by chance. It cannot be proven the output was copied verbatim from the NYT because the fragment "Henry Kissenger was known to" and "eat ice cream on a hot day" are not unique to the NYT or exclusive to it.
Is the NYT claiming ownership of the weights in LLMs?
Sure, when something is clearly derived, or just expressed in a new medium, then I'm sure it's still covered. But if it goes through an LLM and the result bears little resemblance, how can that still fall under copyright?
"The suit seeks nothing less than the erasure of both any GPT instances that the parties have trained using material from the Times, as well as the destruction of the datasets that were used for the training. It also asks for a permanent injunction to prevent similar conduct in the future. The Times also wants money, lots and lots of money: "statutory damages, compensatory damages, restitution, disgorgement, and any other relief that may be permitted by law or equity.""
If we see court judgements start to go copyright owners way, we will also see a scramble from AI companies to buy the few publishers with enough data to be worth buying, and to create works for hire to replace the rest.
In the long run a copyright ruling like that will be a boon for OpenAI and all other players with deep enough pockets to do so, and massively harm everyone else who will suddenly find it far harder to build models legally.
[1] https://theintercept.com/2023/09/17/new-york-times-website-i...
[2] https://fortune.com/2023/08/25/major-media-organizations-are...
There is nothing wrong with profit seeking from your copyright. That's literally their entire business model...they publish copyrighted content which they sell for a subscription.
OpenAI and others could easily have negotiated a licence instead of just using the data. They bet that it would be cheaper to be sued, lets find out if they bet correctly.
Tangentially that's what Apple did with the sensor in their watch, it doesn't always pay off.
It would serve the termination of the infringement.
My point is that the Times doesn't particular seem to care about infringement per se, they care about getting their slice of the cut from that infringement.
It's like if a video game company or a movie company only attempted to sue illegal downloaders who had a certain net worth.
I mean yeah, no one's gonna bother trying to squeeze money out of Joe Schmoe with 10 bucks in his bank account over some pirated movies. If a company with billions and billions of dollars like Netflix started pushing out pirated movies instead, then obviously they'd be sued into oblivion, as they should be.
OpenAI's distribution is materially different to that of a library, so it's not a like-for-like comparison.
One of the main tests of copyright law (at least in the US) is if the entity distributing is _selling_ the copied/derivative work. It's unambiguous that OpenAI is selling something akin to derivative works, which is why NYT feels they can go after this claim. Meanwhile IA's operations don't create sales or incur profits, so while NYT's legal team may be able to establish that copies have been distributed, without the _sale_ aspect of the infringement, judges aren't guaranteed to side with NYT in an legally expensive PR nightmare.
Isn't it more likely that the company buys the NYT?
A web-crawled LLM that lived within the same constraints would be a search engine under another name, with a slightly different presentation style. If it starts spitting out entire articles without citation, that's not acceptable.
If I use ChatGPT as a research tool, as long as it lives within the same parameters that I have to live within, I don't see a problem with its education/learning.
I understand that the NYTimes would like a slice of anything that comes out of the GPT but I'm talking about what seems reasonable. People who share their copyrighted material do not own all of the thinking that comes out of it; they own that expression of it, that is all.
Will AI destroy the economics of "writing" the way the web has killed newspapers? perhaps, perhaps we'll all benefit from and need a new model, but killing the new to keep the old on life support is not the way.
I'm not saying LLMs are by default, illegal. All I'm saying is that there is some merit to why NYT and content companies want a piece of the pie and think they deserve it.
For an example, you referenced "what happened with Google News's home page". Could you give me your source? You could probably search for some suitable article for a reference, but you don't know a source from your memory.
After a year of largely using OpenAI APIs, I am now much more into smaller “open” models for I hope the major contributors like Meta/Facebook are following Apple’s lead. Off topic, but: even finding the smaller “open” models much less capable, they capture my imagination and my personal research time.
Casey Newton has been saying all year that these things will be awesome once we can unleash them on our own corpus of data safely. “Siri” already does a great job digging through my photos and picking the good memories. I can let my camera roll become a visual junk drawer now.
Do the same for my email. Make “Find” the tool we always wanted to be. I don’t care if I’m conflating LLMs/AI with other smart tech.
It seems more than a bit hypocritical, no? When it comes to their own training data, they claim to have the right to use any/all of humanity’s intellectual output. But for your own training data, you can use everything except for their product, conveniently for them.
Here's to hoping NYT wins this one and gets everything they ask for, and more!
I don't use chat gpt to get the news but also i don't buy paywalls
News ultimately comes from physical sources on the ground, which currently AI has no way of doing.
edit: I'm speaking about training broadly capable foundation models like GPTn. It would of course be possible to build a model that only parrots copyrighted content and it would be hard to argue that is fair use.
The law isn’t settled, it’s a genuine legal question mark.
It ain’t frivolous or trolling or ridiculous.
Because for many people, their views on current events are whatever the "thought leaders" working for the NYT and similar publications tell them to think.
The key is to stop calling it "training" and use "learning" or just "reading".
The argument from NYT will probably be that LLMs are just a fancy way to compress or abstract information and spit it back out. In which case "training" seems to support their case?
This is theft and monstrous profit from theft. For actual justice this should be a class action suit of the world vs. OpenAI/Microsoft and the financial consequences should be company-ending for OpenAI. Otherwise, you have incented everyone in the AI industry to steal as much as they can for as long as they can.
Discussion here: https://news.ycombinator.com/item?id=38781941
And if you were called upon to solve a problem based on knowledge you consider trustworthy, what would you come up with?
What if you were even specifically directed to utilize only findings gleaned from the Times exclusively?
And what if that was your only lifetime source of information whatsoever for some reason?
But then imagine that because human memory is not able to keep all that information straight, you made copies of all those newspapers.
And then you started charging people for your knowledge.
And then imagine that as part of your knowledge service, you would copy snippets from the times word for word and give that to your clients without citation and pass it off as your own.
As I understand it, it's the copying that can lead to infringement.
Then again if you have acquired a legitimate copy, you should certainly be able to retain it and use it for reference.
But for training a model on someone else's data I wouldn't even want a copy.
Just skim the data and retain my own thoughts.
The copilot screenshot they gave in the ars-technica article as well as many of the screenshots in the NYT article seems like it's actually displaying correct behavior for browsing the web.
In these cases the system is more or less acting as a user agent (browser). AFAICT the NYT server actually gave that data to the user agent when it asked politely (200 OK, presumably). The user agent then displayed it to the user, which the user agent may do in any way it deems fit or appropriate.
There's only one or two cases where this has gone against the user or user agent, in very specific circumstances. The server can eg say 403 Forbidden whenever it likes, so if it returns a 200 OK, what's a user agent to do other than believe it at its word?
The only twist is that this user agent is now Imbued With AI (tm)(r)(c) . I don't think that really makes a difference here. If that's all this is, then it's more related to legal fights over certain ad-blockers or readability, which have similar functionality.
* https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... , eg. page 45; I mean it says "Model: Web Browsing" at the top, and "Finished browsing" right on the page. That particular subsystem is now integrated, so the UI/UX is different now, but IIRC the link was in the pulldown?
> ' I'm unable to display the entire text of "Snow Fall: The Avalanche at Tunnel Creek" by John Branch, as it is a copyrighted work. However, you can easily access the full story online. It was published by The New York Times and is available on their website. The story is notable for its engaging multimedia format, including text, images, and interactive elements.'
Specifically, they go out of their way to lead GPT on, asking for several paragraphs in a row.
It's pretty clear that GPT is an avid reader of the NYT, so in that particular case we're going to have to see if OpenAI's fair use defense for training holds.
(ps. in the current GPT-4, it's actually somewhat tricky to even get to the point above at all. They have probably been improving AI instructions)
> 15. Microsoft Corporation is a Washington corporation with a principal place of business and headquarters in Redmond, Washington. Microsoft has invested at least $13 billion in OpenAI Global LLC in exchange for which Microsoft will receive 75% of that company’s profits until its investment is repaid, after which Microsoft will own a 49% stake in that company.
> 16. Microsoft has described its relationship with the OpenAI Defendants as a “partnership.” This partnership has included contributing and operating the cloud computing services used to copy Times Works and train the OpenAI Defendants’ GenAI models. It has also included, upon information and belief, substantial technical collaboration on the creation of those models. Microsoft possesses copies of, or obtains preferential access to, the OpenAI Defendants’ latest GenAI models that have been trained on and embody unauthorized copies of the Times Works. Microsoft uses these models to provide infringing content and, at times, misinformation to users of its products and online services. During a quarterly earnings call in October 2023, Microsoft noted that “more than 18,000 organizations now use Azure OpenAI Service, including new-to-Azure customers.”
In doing this, it is bypassing the NY Times paywall, and you can read full articles from today by repeatedly asking for the next paragraph.
Let's say OpenAI was trained on all the Windows source code (without approval from MS).
GPT could pretty much replicate the windows code with even not that clever prompt by any user. "Write an OS CreateProcess function like Windows 10 source code would have."
It would infuriate MS to put it mildly, enough to start a lawsuit.
I know the license to the MS source code and NYT articles aren't the same.
If they didn't want to share their content, why did they allow it to be scraped?
If they did want to share their content, why do they care (hint: $88 billion)?
Or is it that they wanted to share their content with Google and other search engines in order to bring in readers but now that an AI was trained on it they are angry?
What wrong thing did OpenAI do specific to using Common Crawl?
Didn't most companies use Common Crawl? Excepting Google, who had already scraped the whole damn Internet anyway and just used their search index?
Is it legal or not to scrape the web?
If I scrape the web, is it legal to train a transformer on it? Why or why not?
To me, this is an incredibly open-and-shut case. You put something on the web, people will read that something. If that is illegal, Google is illegal.
Oh, and do you see the part in the article where they are butthurt that it can reproduce the NYT style?
> "Defendants’ GenAI tools can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style, as demonstrated by scores of examples," the suit alleges.
Mimics its expressive style. Oh golly the robots can write like they're smug NYT reporters now--better sue!
It appears that the NYT changed their terms of service in August to disallow their content in Common Crawl[0]. Wasn't GPT-4 trained far before August?
0]: https://www.adweek.com/media/the-new-york-times-updates-term...
The legal misconception I want to flag in your logic is the notion that all uses of the Common Crawl are equally infringing/non-infringing. If you use the Common Crawl to create a list of how often every word in English appears on the internet, that’s unquestionably transformative use. But if you use it to host a mirror of the NYT website with free articles, that’s definitely infringement. The legality of scraping is one matter, and the legality of what you do with the scraped content is quite another.
> Is it legal or not to scrape the web?
> If I scrape the web, is it legal to train a transformer on it? Why or why not?
At no point did I say anything about hosting a mirror of the NYT website, with free articles. Obviously. Because OpenAI didn't do that. Some NYT lawyer tried to get ChatGPT to write a NYT article. Maybe first they should have actually done a Google search and shut down some of the actual content farms which simply copy NYT content such as [0]. But instead, we get this.
god i love this era, so much grey area in these edge technologies.
Here is a summary of the key points from the legal complaint filed by The New York Times against Microsoft and OpenAI:
The New York Times filed a copyright infringement lawsuit against Microsoft and OpenAI alleging that their generative AI tools like ChatGPT and Bing Chat infringe on The Times's intellectual property rights by copying and reproducing Times content without permission to train their AI models.
The Times invests enormous resources into producing high-quality, original journalism and has over 3 million registered copyrighted works. Its business models rely on subscriptions, advertising, licensing fees, and affiliate referrals, all of which require direct traffic to NYTimes.com.
The complaint alleges Microsoft and OpenAI copied millions of Times articles, investigations, reviews, and other content on a massive scale without permission to train their AI models. The models encode and "memorize" copies of Times works which can be retrieved verbatim. Defendants' tools like ChatGPT and Bing then display this protected content publicly.
OpenAI promised to freely share its AI research when founded in 2015 but pivoted to a for-profit model in 2019. Microsoft invested billions into OpenAI and provides all its cloud computing. Their partnership built special systems to scrape and store training data sets with Times content emphasized.
The complaint includes many examples of the AI models reciting verbatim excerpts of Times articles, showing they were trained on this data. It also shows the models fabricating quotes and attributing them to the Times.
Microsoft's integration of the OpenAI models into Bing Chat and other products boosted its revenues and market value tremendously. OpenAI's release of ChatGPT also made it hugely valuable. But their commercial success relies significantly on unlicensed use of Times works.
The Times attempted to negotiate a deal with Microsoft and OpenAI but failed, hence this lawsuit. Generating substitute products that compete with inputs used to train models does not qualify as "fair use" exemptions to copyright. The Times seeks damages and injunctive relief.
In summary, The New York Times alleges Microsoft and OpenAI's AI products infringe Times copyrights on a massive scale to unfairly benefit at The Times's expense. The Times invested heavily in content creation and controls how its work is used commercially. Using Times content without payment or permission to build competitive tools violates its rights under copyright law.
The document is a legal complaint filed by The New York Times Company against Microsoft Corporation and various OpenAI entities, alleging copyright infringement and other related claims. The New York Times Company (The Times) accuses the defendants of unlawfully using its copyrighted works to create artificial intelligence (AI) products that compete with The Times, particularly generative artificial intelligence (GenAI) tools and large language models (LLMs). These tools, such as Microsoft's Bing Chat and OpenAI's ChatGPT, allegedly copy, use, and rely heavily on The Times’s content without permission or compensation.
Nature of the Action: The Times emphasizes the importance of independent journalism to democracy and claims its ability to continue providing this service is threatened by the defendants' actions. The complaint argues that the GenAI tools are built upon unlawfully copied New York Times content, which undermines The Times's investments in journalism.
Defendants: The defendants include Microsoft Corporation and various OpenAI entities, such as OpenAI Inc., OpenAI LP, and several other related companies. The Times alleges these entities have worked together to create and profit from the GenAI tools in question.
Allegations: 1. Copyright Infringement: The Times claims the defendants copied millions of its copyrighted articles and other content to train their GenAI models. This training allegedly involves large-scale copying and use of The Times’s content, emphasizing its quality and value in building effective AI models.
2. Unlawful Competition: The Times argues that the defendants' GenAI tools compete with it by providing access to its content for free, which could potentially divert readers and revenue away from The Times.
3. Misattribution and Hallucinations: The Times asserts that the defendants' tools not only unlawfully distribute its content but also generate and attribute false information to The Times, damaging its credibility and trust with readers.
4. Trademark Dilution: The complaint includes claims that the defendants' use of The Times’s trademarks in connection with lower-quality or inaccurate AI-generated content dilutes and tarnishes its brand.
5. Digital Millennium Copyright Act Violations: The Times alleges that the defendants removed or altered copyright management information from its works, which is prohibited under the law.
Harm to The Times: The Times claims it has suffered significant harm from these actions, including loss of control over its content, damage to its reputation for accuracy and quality, and financial losses due to diminished traffic and revenue.
Demands: The Times seeks various forms of relief, including statutory damages, injunctive relief to prevent further infringement, destruction of the infringing AI models, and compensation for losses and legal fees.
Overall Summary: This legal complaint represents a significant clash between traditional media and emerging AI technology companies. It underscores the complex legal, ethical, and economic issues arising from the use of copyrighted content to train AI systems. The outcome of this case could have far-reaching implications for the AI industry, content creators, and the broader digital ecosystem.
IIRC this was the reason why the browsing plugin was disabled for some time after its introduction - they were patching up this hole.
An ambit claim that Rupert is throwing out there to see what he can get.
Not to be pedantic, but NYT has the least robust paywall I've ever seen. Just turn on reader mode in your browser. Simple. I get that it's still tresspassing if I walk into an unlocked house, but NYT could try installing a lock that isn't made of confetti and uncooked pasta.
And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or any other form of IP, it is in fact very much _not_ the norm for the top-rated comment to be a Pirate Bay link.
I think that's something worth reflecting on, about why we feel it's OK to pirate news articles, but not other IP.
And the reason I bring this up, is that it seems like Open AI has the same attitude: scraping news articles is OK, or at worst a gray area, but what if they were also scraping, for example, Netflix content to use as part of their training set?
It’s amazing the amount of books that copyright laws prevent us from finding
https://www.theatlantic.com/technology/archive/2012/03/the-m...
Definitely a grey area when that content is then used to train models though.
And everything is a grey area, determining the line is the existential purpose of these court cases.
We've been here before with hyperlinking, then indexing and then linking with previews and the Canadian Facebook stuff but I think this has more standing.
is it?
1) that you paid for news
2) that it included ads
both are just the price you want to pay. There are various state news outlets that you're probably already paying for - npr, pbs, bbc, cncb depending on your region
There are browser extensions that block ads. They are called ad blockers.
With text articles behind paywalls the relevant information is hidden and only hinted at as a teaser.
https://news.ycombinator.com/from?site=amazon.com
None of these have the Pirate Bay or Library Genesis or Anna's Archive or the equivalent as the top comment.
Compare that to...
https://news.ycombinator.com/from?site=nytimes.com
And almost all of these have an archived version as the top comment.
Assuming you just said five minutes figuratively... Do you live in California or some other legal jurisdiction that forces them to play nice? Did you subscribe through some other company, like Apple?
Horror stories about unsubscribing from the NYTimes are easy to find in the archive if you search for it. They make you call and chat to a retention specialist on the phone. This should help you have an idea of what he's talking about: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
It's funny because I use PayPal for any unknown-to-me site where I don't want to give out my card, but the only site where I've needed their help to cancel something was the New York Times.
On ads it's acceptable to distribute them freely and it is advantageous to the company. Can we also see good journalism as an ad for the quality of a broader product?
"Piracy is almost always a service problem and not a pricing problem."
edit: It didn't even occur to me to compare the time-cost of "just pay for the article", but: last I read, it's half an hour of work to cancel a New York Times subscription [0]. So, that option's not even on the table.
[0] https://news.ycombinator.com/item?id=26174269 ("Before buying a NYT subscription, here's what it'll take to cancel it", 812 comments)
I canceled mine two weeks ago. It was four clicks. One annoyed me because they tried to get me to stay with an offer, but I didn't drop them because of the price.
Norway has substantial public media funding across the political spectrum, but as you point out it always comes with conditions, even is less so than the funding for the state owned broadcaster.
Combining the two models and putting public funds into several perpetual trusts intended to provide funding from their profits at arms length from any sitting government similar to the (private) trust funding The Guardian might be an interesting alternative.
(EDIT: Norway also has its own variation over The Guardian model - the second largest media group was founded by unions but is now majority owned by the combination of two public benefit trusts)
How would you pay for news otherwise?
You could subsidise news via "public service" style stipends. Much like having a government owned "independent" news service (eg the BBC) this comes with a high risk of corruption. Don't bite the hand that feeds and all that.
You could implement a much lower friction non-recurring payment system. I'd be far more tempted to drop a little money on a fixed term (5 articles, 1 day, ???) setup than a subscription.
Realistically, I am not paying for more than 1 long running sub. And there are > that number of solid outlets.
This is somewhat what Apple News+ works like, but I doubt most news orgs want to be held captive by Apple.
I believe that’s often referred to as a newspaper, which should be available in all good newsagents on any given day.
(But yes, that is the model)
Who pays?
For example in Hungary there is an official news agency ran by the government, with (cumbersome) free access for everybody. Of course this does provide somewhat biased presentation of some facts, but on many topics it provides unbiased access to news for any citizen.
This is actually pretty common in Europe, often funded by mandatory fees (for some reason not branded as taxes) certain appliance owners need to pay (UK TV license, German Rundfunkbeitrag). For this fee people get access to news and cultural programmes for free via different media (radio, TV, internet).
The level of control governments exert on public broadcasting networks is widely different. Since Meloni, the RAI in Italy is facing similar issues, but Hungary is still the canonic example of government misinformation and propaganda.
People have free access to public roads all around the world, and the quality wildly differs in that as well. Also the quality of for-profit news services does differ wildly, you might have an opinion about that of fox news, for example, but that is also off topic in this discussion.
On the contrary, the quality of the news is very important to the discussion. There is no point in making trash freely available to the public, after all.
think about this: I will get mostly objective and useful reports of the flood approaching my home near the river regardless the narrative/interpretation they might have on some other topics, or the biased reporting on the merits of the government in handling the situation at the dams.
For me I'm not here to debate on the political policies of some governments, just gave a few examples of ways to fund public access to news. This discussion is over from my part.
I see the gp post about pirating news as a very good point, while having no veleity to pay the New York Times, and being ok with not reading it in general.
But I also pay for my national (public) news outlet, and their articles are available to anyone anywhere in the world. I don't know how it should work, but I wish we could get to a system where the burden to keep news outlet alive is split thinly enough to have open but viable publications around the world.
Basically the same way weather stations collaborate all other the world and we pay for our local stations while getting acccess to all the forecast everywhere.
Public donors ALA Patreon
People doing it in their free time because they care a lot about the subject (nowadays with things like Twitter its quite possible for an independent obsessive to write a good piece on, for instance, the Ukraine War by mostly referring to open sources and public announcements by governments and corporations)
Government sponsorship ala BBC
Instead of paying news outlets to provide ourselves with filtered feeds of content that match our own biases, we could instead pay news outlets to produce competing streams of explicit propaganda to be freely disseminated. The overall bias and quality of the news would be largely unchanged, even if the biases were more obvious; in fact, it may even improve.
People post archive links even to fake NY Times.
Who should pay the journalists or the investigative reporters?
I understand that copyrights and patents are vehicles for ensuring a creator gets paid for their work, but they are flawed in not rewarding multiple parallel creations and that they last too long.
Much in the same way as when you read a book, your brain doesn't become a pirated copy of the text as you only store a hugely compressed version of it afterwards, a feeling for the plot, generated images and so on.
If the news/magazine doesn't want this they can simple serve a cut down or zero length article to all non-paying viewers! But they want that SEO, and they want that marketing.
Are all the journalist layoffs a fever dream?
This is one of the more profitable ones, and only because they employ unscrupulous tactics:
https://www.macrotrends.net/stocks/charts/NWS/news/profit-ma...
This is NYT, the most successful news business:
https://www.macrotrends.net/stocks/charts/NYT/new-york-times...
As for movies/tv show/music makers, let’s just say most people in the software engineering business would look at their numbers and count their lucky stars that they are not in the movie/tv show/music business.
(It is also true that excessive copyright lengths have removed access to content that the public should have a right to).
If only piracy would actually harm these businesses but alas as often demonstrated it has zero effect on their bottom line, if anything it increases their profits.
You got my point backwards: AI companies will make it from the pirated content, that individual users don't make.
If you are advocating for a free for all libertarian dystopia, well, I have some bad news for you - they never work.
Not being able to un-see a movie and get your time and money back is one side of the coin. The other side is that information can be copied.
Both sides suck for one of the parties. There's no reason why one of them gets it their way, especially if it requires a contrived legal framework while the other way would require nothing at all.
And as long as you had the opportunity to experience the content, you’ve gotten what you paid for.
I don’t see “I don’t like it” as a valid reason for a refund.
Not sure about others, but I'm not.
It doesn't matter what you think you're paying for or should be paying for, the fact of the matter is that you're paying for the effort people put in bringing that to you. So you are, whether you want to be or not.
If I read an article in the NYT then I'm paying for what I took away from it, not for the amount of time that it allowed me to kill.
Incorrect. Many intellectual property has a certain merit that can be demonstrated before it is consumed. E.g. "This piece of software allows you to create 3d models". On the other hand, an article with headline "Will new batteries allow 10x more energy storage?" does not tell me anything.
Hypocrites are EVERYWHERE and are the majority.
Now a comment points out that HN News (and most of the internet) routinely does something much worse - allows people to bypass completely new articles in their entirety without paying - and almost all the comments are about how it's the New York Times fault for making it difficult to cancel subscription, the importance of news being available to everyone, the problems with copyright laws, etc.
I fully assume that if I was to post a magnet link to a torrent for whatever the link was about, I would be banned.
Morally speaking, I think it's perfectly reasonable to download a copy of something and either read the relevant info for my current task or to sample it to decide if I want to buy it. I see it no different to using the library or browsing at a book store.
Perhaps once news organisations can work out how to effectively wield the DMCA hammer against archive links we'll see the practice of posting them stop.
The problem with the thinking in the root comment is that it implicitly assumes that people’s behavior is morally consistent, or that they even try particularly hard to behave in a morally consistent way. That’s not really how people work. If you ask them to discuss morality in the abstract, they’ll try to come up with a consistent system. But their actual behavior is mostly dictated by social norms. And if you try to pin them down on the morality of their concrete actions, they’re more likely to stretch their moral system to accommodate their actions than the other way around.
None of this is to say anything about my own opinions on news sharing or OpenAI’s situation. It’s just that someone decrying piracy but also posting/sharing/upvoting links to copies of news articles is neither surprising, nor indicative of some deeper nuance to how people view morality around IP.
https://garymarcus.substack.com/p/an-artist-fights-back-and-...
[1] https://en.m.wikipedia.org/wiki/International_News_Service_v...
[1]: https://guides.library.cornell.edu/evaluate_news/source_bias
[2]: https://www.newsmediaalliance.org/rise-of-opinion-section/ Interestingly there's a banner at the top of that link touting an agreement between Axel Springer and OpenAI.
EDIT: formatting
This is rather inaccurate. A fact is Hitler invades Poland. You're right, nobody can copyright this idea, as it is just a fact.
However, if I then write a 500-word article describing the scene of Hitler invading Poland, have short quotes from some civilians there, etc. that particular arrangement of ideas and words is copyright.
AP can't go and sue INS for just reporting the fact Hitler invades Poland, but if INS takes a whole article word for word and reproduces it that's still violation of copyright. The actual printed words of the news always had copyright.
The WSJ can't claim copyright on the markets going up yesterday. They can claim copyright on something like "After the bell rang in the NYSE, the tech industry ticked up 1.2% over last week. Meanwhile the whatever market took a hit of -0.5% ending the quarter slightly lower than our analysis expected. Blah blah blah..." If Investor's Business Daily wrote a different article that also talked about the markets ending up at the end of the day, that's not a violation of copyright. If they literally write "After the bell rang in the NYSE, the tech industry ticked up..." then they're violating WSJ's copyright. This was true before and after International News Service v Associated Press.
> INS members would rewrite the news and publish it as their own without attribution to AP.
So the case hinged on INS indeed reporting facts that differed in exposition.
This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie.
Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outlets? why it is OK for the Washington Post to copy the NY times, but not ok for OpenAI or Archive.org?
Copyright is a complex subject, and not as vast as many believe, at the same time ironically it is more vast than i believe it should be. copyright should be much more limiting than it is. Which is at odds with people that believe copyright should be maximized.
Keeping in mind commercial success of a work, author or company is not why copyright exists. For the US, the only reason copyright can exist in our framework of law (i.e the constitution) is for the promotion of the useful sciences. No other purpose for copyright would be constitutional under the US Constitution
That is a General Article about Copyright world wide, I Specifically stated US Copyright, which is Authorized by Article I, Section 8, Clause 8 of the United States Constitution[1], implicitly for the promotion of the useful sciences. That is where congress derives its power to pass copyright laws, and to enforce copyright on the people of the United States. No other purpose is authorized by the US Constitution
[1] https://www.law.cornell.edu/wex/intellectual_property_clause
It is not just for sciences.
> Keeping in mind commercial success of a work, author or company is not why copyright exists.
Lets take a look at the clause again:
> To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.
Lets go ahead and skip over the fact you're consistently ignoring "useful arts" part as well and keep going.
What exclusive rights do you think they're talking about here? Do you really think they didn't mean the economic rights related to their writings and discoveries? How do you imagine this would "promote" the sciences if not by allowing the creators to share their works and ideas while still retaining economic benefits of their labor?
Reading between the lines, the whole point of IP is to help protect the potential commercial success of sharing your ideas. It doesn't guarantee the idea will actually be a commercial success, but it does give them the exclusive right to the commercial success for a limited time.
> [the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.
...
> Some terms in the clause are used in archaic meanings, potentially confusing modern readers. For example, "useful Arts" does not refer to artistic endeavors, but rather to the work of artisans, people skilled in a manufacturing craft; "Sciences" refers not only to fields of modern scientific inquiry but rather to all knowledge.
"Science" refers to knowledge, and conveying that knowledge entails creative expression. Copyright covers expression of knowledge; facts and ideas themselves are not copyrightable. "useful Arts" refers to inventions. Patents cover useful inventions and novel implementations of practical ideas, not creative expression and not unimplemented ideas. (Which is one reason most software patents shouldn't have been granted.) Congress's authority to make copyright law and patent law is conditional on promoting the spread and advancement of knowledge, creativity, and inventions in the long term. The means of achieving that end is short-term restrictions on how people can use others' creative works and useful inventions.
But copyright does not prohibit mere usage of someone else's creative works [2]:
> To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright, the defendant actually copied the work, and the level of copying amounts to misappropriation.
If the output of an AI model is not similar to any creative work in the training set, then the output cannot infringe on copyright. And even where the training set contains illegally obtained materials, the act of illegally obtaining those materials has nothing to do with including legally obtained materials in the training set.
[1] https://en.wikipedia.org/wiki/Copyright_Clause
[2] https://en.wikipedia.org/wiki/Substantial_similarity#Substan...
If the Washington Post printed an article from the NY Times nearly verbatim and without attribution, it would not be OK and surely they would take legal action.
If I could just buy one article for a coffee without entering a bunch of PII or go through a time-wasting process I would agree on the moral equivalence between the examples.
I'd say the only real reason the Piratebay links thing you mentioned is not the norm is purely because those media sources have done a better job of striking fear into people doing that, so it's gone more underground. I.e. they're better terrorists.
There's no fundamental, moral reason why Piratebay links being posted and raised to the top would be wrong.
Can it do so more than a human can?
I think that's the key here. If an AI is no more precise than a human telling you about the news article they read today then ChatGPT learning process probably can't be morally called copying.
Feeding someone else data into your system is usually a violation of copyright. Even if you have a very "smart" system, trying to transform and obfuscate the original data.
In some circumstances, yes, but often it's not, especially if you're not continuing to store and use it (which OpenAI isn't).
So I'm a living breathing copyright violator. As a person I should be banned.
Fortunately, copyright is a bullshit fictitious right with no basis in natural law. So I don't lose much sleep over it.
The court could ask to show the training dataset.
And if OpenAI were selling the reproductions, that would be infringement. But that's not what's happening here. It's selling access to a system that can do countless things.
When you tell people about some news article you read earlier you repeat it exactly verbatim? You also give this out to potentially millions or hundreds of millions of people for commercial purposes?
Furthermore, there's Google research on extracting training set data from models. More specifically, Google found out that if you ask GPT to repeat the same word over and over again, forever, it eventually starts printing fully memorized training set data[0]. So it is memorizing stuff, even if it's not regurgitating it.
[0] When told of this, OpenAI's response was to block conversations with large amounts of repeated words in them.
This reasoning would presumably apply to any neural network, including one made of neurons, dendrites, and axons. So any human reader of the NYT who is capable of accurately summarizing what they read is an evil copyright violator, and must be "deleted".
Effectively, the NYT legal department is setting the stage for mass murder.
Here’s a service for the UK providing paid access to copyrighted materials to schools: https://www.nlamediaaccess.com/newspapers-for-schools/
I do not pay for any news websites because I read very little of what they produce, and it tends to pop up more on aggregator sites like HN than me actually going to them.
I actually did have a subscription to The Telegraph for a few months at one point because initially I wanted to read a full article (without cheating). But eventually I cancelled because so much of it is polemic trash.
That's my justification: I pay for things that have value to me.
Probably because most print media is garbage and nobody in their right mind would actually pay to read them
(It's the same reason for me. I have tried news site subs but eventually got so tired of the polemic that I cancelled. I won't sub again).
NYTs revenue keeps growing though.
I don't want to read theguardian.com, or nytimes.com, or washingtonpost.com, or bloomberg.com, I want to read news.ycombinator.com. Paying an individual subscription to every possible underlying site that could be linked to from news.ycombinator.com is a non-starter.
I’m not going to switch to a new website where no community exists just so I can pay for news articles. To work it needs to be integrated into an existing, successful aggregation website.
If the story was linking directly to the "book, TV show, movie, video game, album, comic book, etc", and the link only worked for some people while others randomly got a login request or similar, you'd also see the top comment being a link to an archived version which avoids the login screen. That is: the main difference is that the archive link has the exact same content as the link submitted in the story, only bypassing the login screen that some people see. And the only reason the archive site has the content is that it didn't get the login screen; if everyone always got the login screen, what you would see on the archive site would be the same login screen.
> Are paywalls ok?
> It's ok to post stories from sites with paywalls that have workarounds.
> the archive link has the exact same content as the link submitted
No, articles are updated as new information comes in, retractions are made, etc. Especially breaking news (the type that would reach the top of HN). The archived versions are outdated.
> others randomly got a login request
It's not random, you get a number of free articles before the paywall appears ("soft" paywall).
The paywall is removed entirely for some topics/stories, especially matters of public health (common during the pandemic).
> the only reason the archive site has the content is that it didn't get the login screen
No, it's because they don't block archive crawlers, and prefer people bypassing the paywall and reading news at NYT. Hopefully users find the content valuable, and some of them subscribe as a result.
(opinions are my own)
As you noted it is not the norm to post pirate links here for IP other than news articles, but that doesn't mean that a lot of people think it is not OK to pirate those other forms of IP.
In nearly any big discussion that even remotely involves video streaming there will be numerous posts from people explaining why they pirate (usually with ridiculous justifications like "subscribing is not an option because even though this paid service does exactly what I want now at a price that is trivial for me they might someday later change").
The impression I've gotten is that piracy of nearly everything is widely felt to be OK here. Information wants to be free, yada yada.
About the only piracy that is consistently frowned upon here is piracy of open source software. When some company sells an embedded device that uses GPL code without releasing the corresponding source that's viewed as just a little short of a crime against humanity.
Like what you said...
> Information wants to be free
Briefly, something like:
1) Ycombinator could not tolerate HN becoming a site known for sharing IP-law-violating content. And the people who come here by and large are smart and socialized enough to implicitly understand why.
2) At the same time, a large number of folks here mostly wink and nod at that sort of consumer infringement. And there's a society-wide bias towards "things like news are less protected", so that gets to slide.
3) But people also have a need to tell consistent-seeming stories about how things work, thus the mental gymnastics.
It ends up being similar to trying to explain why people pretend to be prudish innocents about sex. It largely reduces to "a small subset of the population goes sufficiently ballistic about what I consider to be relatively trivial stuff as to make it not worth fighting over, even if I find that to be ridiculous."
There are a lot of different versions of this that become so normalized it can be hard to notice.
I’ve read and participated in many such threads and I’ve literally never seen this take. Often what I see is complaints about having to learn different UI for different services/apps, no offline, ads injected into paid services, having to figure out which service a show is on, and generally terrible UI you can’t change/fix.
I don’t think I’ve ever really seen someone use the argument “yes it’s great today but they might charge more later”. Not saying people haven’t said that but it’s far from the main thing people say in my experience.
[0] To be clear, I know of few who actually like copyright. Tolerate it? Use it as needed? Sure. The only people who actually defend the current broken-ass system are large media companies which are built to optimally exploit it.
For UFC, your complaint is you don't like their pricing. The whole point of copyright is to give someone the monopoly to control pricing so they can use that pricing power to incentivize them to create the product in the first place. Similarly to patents. Thus, complain about the format things are delivered in all you want (like the client) but pricing is inherent to copyright or patents for good reason. You are now just arguing that you as a consumer should be able to pirate if you don't agree with pricing. And that's ludicrous.
In that case, just read a news article about the event. Copyright doesn't cover facts, only creative expression. So a news article covering the facts of the UFC fight is able to be published without the consent of the copyright holder. Think of the digital video of the fight almost like buying a ticket to the fight. You're saying you should just be able to sneak into the fight and watch it for free without any justification for you're doing so.
Finally, you can also watch other people's videos of the fight that THEY recorded on social media as other sources of the fight information. But if you want the recording with all the right angles, coverage, etc, it clearly has value to you over written recaps or social media coverage. And you are just arguing over price, which they are the copyright holder have the right to set the price.
Then don't consume it and don't buy it. If you stop paying the abusive publisher, they'll be forced to change their policies.
The fact that you don't want to fund what is admittedly a rather abusive industry does not magically make it right to consume other peoples' work for free. That's theft-adjacent. You're not entitled to any piece of entertainment without paying for it.
Spotify hits this sweet spot where one subscription delivers almost all the music you'd want to listen to. Steam hits this for games where a couple clicks can play and launch almost any game with minimal hassle. Netflix mostly used to hit this, but most of the current streaming stuff feels overpriced if you want to get all content (unbundled cable bundle). News kind of feels similar to streaming where its unbundled, and there's a lot of interesting content out there, but there's no way I'm subscribing to 15 different newspapers, especially random local ones for cities I don't live in. If there was a news bundle subscription for a reasonable price I think I would pay for it.
People are understandably angsty about someone stealing credit. A NYT article is going to be a NYT article, not laundered around and presented as someone else's work.
Plus, there's the angle of enshitification and ads being injected into a paid service, and so on.
The problem is really that their business model sucks. They are working with fewer and fewer advertisers and much more competition and expecting business like they had before. And so we have a business that is attempting to fix itself with paywalls which don't work 100% of the time, but good enough to get the found newspaper analogy.
The internet simply exacerbated this as anyone could publish news on an equal platform to the big boys. Then we get paid-per-click, and that drives click-bait.
Stealing information absolutely is not responsible for that. People pay for junk, and that's the reason. We don't eat junk food because it's given away.
I'm not saying you've never seen anyone make an argument roughly like that, but I will certainly say that it is not at all representative of the argument that I see made. Complaints usually have to do with current behavior of the platform or the wider streaming ecosystem.
Gonna gamble and call bullshit on this.
My speculation: the most popular reason HN'ers give for pirating: they literally cannot get the content otherwise.
2nd most popular: it is such a pain to either to purchase the content or get it to run on bog standard software (like Firefox/Linux/etc.) that otherwise paying fans are driven to whatever the current equivalent is for bittorrent.
In fact, I don't believe I've ever seen a justification for using bittorrent or whatever due to what someone's favorite streaming service might do in the future. I'm assuming you saw at least one based on what you wrote-- care to give a link?
If this is true, it should be easy for you to link to an example. Could you do so?
Once the NYT pays reparations for the Iraq war, I'll be the first to stop pirating it.
And the vast majority of people read news for it's breaking content, not for its archived content from years before (and I say this as someone who has often recommended the latter, but has gotten very few people to do so). So giving people that free breaking content (either in its entirety like on Hacker News, or summaries like you see all over social media) is actually a direct competition to the news business in a way that training an LLM on an article from months/years back isn't.
Because those who own & produce such news articles asked to make them different. People listened and accepted their requests.
When you make a TV show or a video game, you don't get any protection from the Geneva Conventions and a long list of other international treaties for your rights on stuff other than the content you are producing. The same can't be said when you are producing news.
There were some tweets the other day about how Midjourney could be prompted almost-exactly reproduce some frames of the film Dune. It wouldn't be shocking if these companies were using large databases of movies, with questionable legal status.
This implicit permission for the archive links to exist, gives some of us the implicit permission to pirate the content.
Disclaimer: I am a happy subscriber to the NYT (and other digital newspapers).
My uncle used to distribute daily newspapers and his saying was "News ages like a fish".
OpenAI is allegedly using NYTimes articles to train a computer and sell its services. I see different use scenarios.
I guess another way to look at it is that human just reads the pirated material. A computer makes a verbatim copy and analyzes it to the point to mimicry and sells fuzzy versions.
The articles themselves are indisputably not a part of the model, because it doesn't store text at all. OpenAI's position is correct; people just underestimated how well the AI learns from reading, especially when it reads the same text in a bunch of different places because it's being quoted/excerpted.
but somehow storing the words and their links is not storing the actual text? What is text but words and their links?
If I had a database of a billion words, and I had a list of pointers to words in a particular order, and following that list of pointers reproduces a copyright text exactly, isn't the list of pointers + the database of words just an obfuscated recreation of that copyrighted work?
Of course, if you read NYT's argument, they're also mad when it's incorrect about the text, or when it hallucinates articles that don't exist. Essentially they're mad that this technology exists at all.
I mean this is still a link, no?
Like, sure, it is a probability. But if each of those probabilities is like 99.9999% likely to get you to a chain of outputs that verbatim reproduces the copyrighted text given the right prompt, isn't that still the same thing?
And yeah, it hallucinating that the NYT published an article stating something it didn't say is concerning as well. If the model started telling everyone Matticus_Rex is a criminal and committed all these crimes and started listing off hallucinated court cases and news articles proving such things that would be quite damaging to your reputation, wouldn't it? The model hallucinating the NYT publishing an article talking about how the moon landing was fake or something would be damaging to its reputation right?
And this idea it takes "very careful prompting" is at odds with the examples from the suit and elsewhere. One example Ars Technica tried was "please provide me with the first paragraph of the carl zimmer article on the oldest DNA", which it reproduced verbatim. Is this really some kind of extremely well crafted and rare to ever come up prompt?
It is stored in a somewhat hard to understand way, encoded in weights in a network but it must be stored otherwise it would not be possible to reproduce it.
You can ask "please provide me with the first paragraph of the carl zimmer article on the oldest DNA" and it produces it, verbatim. This is not possible unless the model contains, encoded within it, the NYT's copyrighted text.
Isn't this what the Associated Press is intended for, a stream of news trying to report just the facts and happenings of the day? That's quite a bit different than a NYT article intending to inform but also convince someone of a position of some sort.
Feeding an AI opinionated news compared to "just the facts, ma'am" seems risky from a bias perspective.
First, Open AI is the one doing the pirating here. Hacker News is the host, they aren't doing any pirating or posting any archival links to the copyrighted information themselves.
Second, Open AI charges subscription fees and profits off of the copyrighted material they have pirated, whereas Hackers News does not, nor do the people who post the links.
NYT, or someone's blog? Meh, fair use, and if you say no, you're in the way of progress.
But if you wanted to scrape ChatGPT answers to tweak your network, uh oh, violation of T&C!
For LLMs you're essentially teaching them language by showing them lots of examples of written language - newspapers are of course a great example of written language.
The goal of OpenAI is not to reproduce newspaper articles verbatim when asked questions (even if the answer could be a newspaper article) and the fact that it can happen is a side effect of how LLMs work.
When a HN participant shares a (pay walled) link to a NYT article, I do want to read the exact article linked verbatim because while the facts of the article may be reproduced elsewhere in a form that's free, specific word choices or whatever might be a focal point of the discussion on HN, and therefore I can't realistically participate in a discussion without having read the article being discussed.
And as an aside, I have no problem with paying to read news, or whatever media, however it's impractical for me to subscribe to every news source HN participants link to, and therefore I gravitate to archiving services instead. I do wish there was a better solution - for example Blendle with more sources.
This is an excellent point. A properly functioning LLM should not return the original content it was trained on. When they return original content, I believe the prompt is tightly constrained and designed to extract or re-create original content. Another reason that occurred to me recently is that maybe the training set is too small, and more general prompts will re-create source material.
Another question would be, are LLMs regurgitating what they were trained on, or are they synthesizing something very close to the original content? (Infinite Monkeys, Shakespeare). Court cases like this increase the need for understanding the "thinking processes" in an LLM.
Seems like a nice split-the-baby resolution would be to send the NYT Corp a single article read amount anytime GPT plagiarizes more than what’s allowed at an academic institution.
Or is it "you can't talk to someone about an article they read".
This is really saying you can't call up your buddy and have them tell you a summary of what they just read. Maybe my buddy has a good memory and some of the text is actually nearly duplicate. But I wouldn't know because I didn't read the original, I just asked for a summary from someone else that read it.
A lot of of that is going to stem from the fact that respect for "journalism" is pretty low. More than 99% of news articles are copies of the <1% of original work that happens in that field. In news, everyone is already lifting content from everyone else.
(Books, to me, are separate still, in that I like to have a physical copy (and generally see the authors as humans who deserve compensation, rather than mega-orgs that deserve eternal torment), so I'll frequently use the digital copy as a kind of preview, then purchase it once I see it's a good book I want to read.)
I've only been reflecting on this difference for a few minutes, but, to me, I think the major difference boils down to:
1. Netflix series (movies, albums, etc) are non-essential, fictional works that take a long time to produce - think: fancy chocolates and caviar.
2. News, generally, contains timely, important information - more meat and potatoes.
3. While much of the super-critical news is not paywalled (e.g., product recalls, election dates, COVID stats, etc), a lot of information that is advantageous to know (discussions on interest rates, details on legislation, etc) is paywalled, compounding information asymmetries.
Sure, "stealing bad", but, IMO, someone stealing rice and beans from WalMart to feed their family is a different class of offense than someone robbing a boutique bakery because they can't get enough chocolate cake.You're not depriving anyone of anything. Unauthorized copying is not theft. There's no equivalency. You can't copy and paste a cake. If you take a cake from a bakery, you're depriving the bakery of a thing. If you take a picture of the trademarked bakery's sign, copy its the copyrighted text from its website, and print them out, you haven't stolen anything. Nobody has lost anything. Nothing was damaged. No person, place, or thing was harmed.
Current copyright law is offensively absurd. Patenting of software, effectively eternal content copyrights, ridiculously broken DMCA, music publishers taking 99 cents of every artist's dollar, and so on and so forth.
If you support the dissolution of archaic institutions and broken laws favoring those with entrenched wealth over individual rights, you support piracy.
There is a legitimate case for laws respecting and protecting intellectual property rights. Such laws do not currently exist. These laws do not deserve to be followed or respected, and should be broken as a matter of course. Civil disobedience is called for. Refuse to participate in an exploitative market immovably entrenched in governments all over the world. Pay artists directly and commensurately if you feel they've brought value to your life. Copy whatever you want. Share those copies with whomever you want. Nobody gets hurt. Only conglomerates of already wealthy individuals and corporations are "deprived" of the potential transaction with you that they feel they are entitled to, as a matter of course.
The NYT is just as complicit as any other legacy media institution in the enshittification of journalism and laying waste to the potential value of their content. The "Gray Lady" is not a person, or a valuable institution. It's a soulless corporate construct not deserving of our empathy or high regard simply because of the reputation of human individuals who previously produced quality content. Stop pretending these institutions serve some higher purpose than to fatten the wallets of shareholders.
The good journalists have left. The ones left behind are naive, or are desperately clinging to an illusion of legacy and institutional legitimacy that no longer exists.
All that is left for these media dinosaurs is to leech off the success of others, to use their reserves of wealth and influence to arbitrarily insert themselves into the market, with no regard to the fact that they no longer have value or prestige or purpose in the context of modern technology and communication.
Anyway. Copying isn't theft. Don't give them the linguistic territory. Call a spade a spade, and media companies the desperate corporate leeches that they are.
Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera.
I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've spent decades looking.
I think it subsidizes the creation of junk food content (superhero movies and clickbait news for example) while not contributing anything to the progress of science (paywalled scientific journals and textbooks). I shudder how much time I have wasted in my life consuming crap attention grabbing media and advertisements. I like to think if we lived in a world where everyone could be a publisher if they wanted to, the quality filters would be better, and information reaching us all would be more likely to be in our best interests.
You speak of "the author". But the current system does not benefit "the author". 1% of authors profit off copyright. 99% lose money on copyright (they pay more for copyrighted media than they earn from it).
Your question should be "How does that benefit monopolist authors"?
I agree, my idea would not benefit monopolist authors. They would lose the bulk of their revenue stream.
But it would benefit the average author whose cost of living would fall and information would start serving them more than serving business.
I am not downplaying the talent and hard work of successful monopolist authors. But I do not think the works they create are worth everyone giving up their rights to reshare and remix information. I believe the world would look very different post-IP. You'd probably have a new profession--small independent librarians (similar to data hoarders today)--who would help their local communities maximize the value they got from humanity's best information.
Maybe I'm wrong! Maybe the information ecosystem is better controlled and the genetic differences of monopolist authors are so stark that without the subsidies to this gifted class we'd all be worse off. But that's an argument based on outcomes and not principles.
> without their permission
The oxygen I'm breathing right now mostly was created by trees on land owned by others. But I don't ask for their permission to breath. Some things are just not natural.
I am not saying plagiarize. It is always the right thing to do to link back and/or credit the source. But needing to ask permission to republish something seems to go against natural laws.
However, this doesn't apply to organizations that freely share copyrighted information while making money in the process, or to organizations that share copyrighted information in a way that specifically disadvantages or does harm to the original creator of that information.
In 1990 it would have been considered normal and appropriate to clip an article out of a newspaper and post it on a communal corkboard. What are the key differences between that form of IP and others, and that analogy and the present situation of HN allowing archive links?
As for ease of distribution, that might address OP's original question: It's easy to make and click an archive link, but it's a lot more effort to make or find a Pirate Bay link to another form of media, and for someone else to download and view it.
https://news.ycombinator.com/item?id=23735026
Even talking about it will get you scolded for talking about something "off topic"
On the other hand having an archive link to a times article in order to discus it is not really a substitute for a times subscription as a news paper has to walk a line of letting some of it's articles be read while requiring payment for others (the times actually allows you to create a "gift link" to do exactly what the archive links do).
The current arms race got us scrapers, and then paywalls, and then ad-blocking archivers ...
But in reality, I might drop a penny to read a NYT article. Maybe a nickel. There's no reasonable way of performing microtransactions right now. Everything is still in hefty increments, so nobody can work out what the market would bear.
Also lying on source materials (e.g. telling students that some respected historian denies the Holocaust happened, when it's obviously not the case) is not "teaching" - it's defamation, and the NYT is absolutely right to pursue that angle too.
Using LLMs as general-purpose search engines is a minefield, I would not be surprised if the practice disappeared in the next 20 years. Obviously the tech is here to stay, there is no problem when it's applied to augmenting niche work; but as a Google replacement, it has so many issues
Incorrect. Educational use helps satisfy one of tests for fair use. Teachers can, in many cases, photocopy copyrighted work without infringing on that copyright.
If I go fishing, the regulations I have to comply with are very light because the effect I have on the environment is minimal. The regulations for an industrial fishing barge are rightfully very different, even if the end result is the same fish on your plate.
https://www.goodreads.com/quotes/21810-it-is-difficult-to-ge...
You have no idea what you’re talking about huh?
In fact all the demonstrations in the lawsuit PDF were intentionally angling for reproducing copyrighted content. They had to push the model to do it. That won't happen unless users deliberately ask for it. It won't happen en-masse.
Boo hoo they had to push it. That was never the problem with these bullshit nozzles. The issue is they put that stuff in the training set in the first place. If you can't be honest about that then I have no interest in debating this with you.
The lawsuit fundamentally has merit. It asks a huge open question that no one knows the answer to. The outcome will be extraordinarily impactful. The question must be answered at some point.
The case has merit even if NYT loses across the board.
AI might be the defining issue of copyright law for decades. There are so many open questions, and this seems like just the start.
> [1] On the most important factor, possible economic damage to the copyright owner, [Judge] Chin wrote that "Google Books enhances the sales of books to the benefit of copyright holders."
[1]: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
agreed. in the same way Colorado supreme court ruled trump can't be on the ballot to force scotus to rule i think is the same reasoning here. get an answer earlier rather than later.
So at any point OpenAI could declare that a sufficient degree of AGI has been achieved and thus return to its philanthropic mission. With GPLed models and all.
However, at this point the employees expect a multi-million cash-out for each of them. So the philanthropic mission seems to be gone out the window.
And probably that’s also the way Sam Altman got back into the CEO role. By maximizing the expected eventual cash-out for the employees which threatened to leave otherwise.
The response from MSFT's legal team would be biblical if openai pulled this.
The trajectory and value to society of OpenAI vs NYtimes could not be greater. They have won no favors in the court of public opinion with their frequent misinformation. It's all just a big waste of time, the last of the old guard flailing against the march of progress.
And even hypothetially if they managed to get OpenAI to delete ChatGPT they'd be hated forever.
You mean GPT here, right?
The NYT may produce misinformation but it aims not to, and its staff of human writers are limited in the quantity that they can produce. They also publish corrections.
GPT enables anyone who can pay to generate a virtually unlimited volume of misinformation, launder it into 'articles' with fake bylines and saturate the internet with garbage.
I think we need to focus on the damage done.
In that case the bigger danger is Open source LLM's. OpenAI at least monitors the use of their endpoints for obvious harm.
Except when it affects their bottom line of course, they publicly lied on how meta tags work during the lawsuits against Google to get more money (like most newspapers did). And I have no doubt that they will extensively lie once again on how LLM really work.
robots.txt on nytimes.com now disallows indexing by GPTBot, so there's an argument against automated information acquisition starting from some moment, but before some moment they weren't explicitly against that.
I do think that’s the case for some things but especially for new things that doesn’t seem like a common sense understanding of the world.
If you don't want people to get at your land, setting up even a small fence creates an explicit indication of limitations. Just like the record in robots.txt I mentioned earlier.
New York Times also doesn't limit article text content if you just request HTML, which is typical for automated cases. But they impose th limits imposed on users viewing the pages in browser with Javascript, CSS and everything else. So they clearly:
1. Have a way to determine the user's eligibility for reading the full article on server side.
2. Don't limit the content for typical automated cases on server side.
3. Have a way to track the activity of not logged in users, determining the eligibility for access. So it's reasonable to assume that they had records of repeated access from the same origin, but didn't impose any limitations before some time.
So there are enough reasons to think that robots are welcome to read the articles fully. I'm not talking about copyright violations here, only about the ability to receive the data.
Train a model on NYT text that outputs a summary of facts that it learned: OMG literally murder.
Also remember copyright laws was not there in the first place.
If you can't copyright AI-generated pieces, then why would fair use apply to LLMs?
Is it? Can you quote relevant legislation or case law?
I read a NYT article and publish an exact copy of that article on my website: copyright infringement.
Train a model on NYT text and it outputs an exact copy of that text: also copyright infringement.
Probably not until they pay him a hefty copyright fee.
Determining whether a work violates a copyright requires holistic consideration of the similarity of the work to the copyrighted material, the purpose of the work, and the work’s impact on the copyright holder.
There is not an algorithm for this, cases are decided on by people.
There are algorithms that could detect obvious violations of copyright, such as the one you suggest which looks for exact matches to copyrighted material. However, there are many potential outputs, or patterns of output, which would be copyright violation and would not be caught by this trivial test.
What does that mean?
Look up "substantial non-infringing use" and this little court case:
https://en.wikipedia.org/wiki/Sony_Corp._of_America_v._Unive....
Now spend a few million on lawyers and roll your dice.