NY times is asking that all LLMs trained on Times data be destroyed
twitter.com
twitter.com
I’d think that the clothing industry went through something similar already.
I no longer sell code, just my services.
I did have one case where someone stole my code, and then tried to sue me for copyright infringement. Too bad for him that I had a registered copyright on it :-)
This will drastically increase the total output of all the "brogrammers" maybe even enough to replace some of the spagetti code out there with more reliable work.
Besides, AI only produces a sort of "average" of what is already out there. Is it really creative work to produce an equivalent of that average?
I remember when the "typing pool" was a thing. Word processors utterly destroyed that category of work, as well as the jobs for typesetting and layout, back in the 1980s.
More generally, AI introduces yet another systemic mechanism for wealth extraction by the wealthy. The wealth here is creative power of the !wealthy. More sinister than the grand larceny by a thousand minor borrowings is the fact that meaning, ideals, motivating energy to move masses is taken away from the candidate pool — anyone of us can be the next demagogue, it’s not too late yet — which may include genuine thought leaders are going to be buried by the electric demagogue working for the proverbial man. “We need to hire more copyrighters for our propaganda using these ‘word processors’” becomes “Who could resist the onslaught of “our” creative efforts? Surrender now Dorothy.”
tldr:
Humanity has reached the absolute limit of the utility of ancient means of governance. New technology demands that we comprehensively review socio-economic order in society. Failure to do so will gift the current “winning” players in the ‘zero-sum-game’ of the du jour regime near guarantee of perpetual habitation in their very very special social perch.
“Think of the children”
NYT and similar publishers are just looking for a new income stream. Rent seeking in the new AI age will probably be more profitable than producing new content.
When LLMs reproduce someone else's work, it is theft and appropriation, and our tech culture is rationalizing it out of hatred of the laptop class -- the kulaks, in my analogy. These traditional media do have their faults, and it's fine to point those out. But reporting is work, consisting of more than just "copying and transforming", and stealing it is wrong. "[F]orcing" it into the public domain is no more a "best outcome" than forcing the farms to collectivize.
I'm going to go out on a limb and say that solution wouldn't be acceptable to the NYT, which is why I think they're trying for a land grab. They're trying to extend copyright beyond what was intended.
I would seriously consider to take it even further. Require that all copyrighted material be made available for public model training.
Who owns the copyright then ?
If you ask for Harry Potter and it gives you Bart Simpson it’s useless.
Technology that makes copyright violations easier/quicker have typically been found legal if "the technology in question had significant non-infringing uses".
They're just on a hunt for some extra money.
I have a gadget that will, with some probability, steal your life's savings. It operates through a process that is analogous to a human chewing. When engineering it, we just say for simplicity that the gadget "chews". Of course, that's only a metaphor -- machines can't chew.
But (and here's where your argument gets ridiculous), unless you can quantify the fact that my gadget can't chew, then I will steal your savings. Good luck.
I can think of 2 instances of that machine already. the finance industry fees and an ex-wife.
1. Training an LLM is akin to human learning. It is legal to read a textbook about music to learn music, and later to write a book about music which likely includes some of the concepts you earlier learned.
2. Neither the LLM nor the output text contain sufficient elements of the copyrighted work to qualify for copyright protection. Just like if you turned old library books into compost and sold the compost, you wouldn't expect to pay authors of those books a royalty for the compost sales.
If you learn a little too hard though, and reproduce the original textbook in it's entirety, you'll get in trouble.
My guess is that courts will determine that the training itself will not be found illegal, but either the AI companies, or the users, will be found liable for reproducing copywrighted work in output, and no one will want to hold liability for that.
AI developers will most likely rely on a Fair Use defense. I think this has a reasonable chance of success since, while the use of a given copyrighted work may affect the market for that work (in this case NYT's article), it can be argued to be highly transformative usage. As in Campbell v. Acuff-Rose Music: "The more transformative the new work, the less will be the significance of other factors", defined as "whether the new work merely 'supersede[s] the objects' of the original creation [...] or instead adds something new".
There's also potential for an "implied license", as in Field v. Google Inc for rehosting a snapshot of a site, where "Google reasonably interpreted absence of meta-tags as permission to present 'Cached' links to the pages of Field's site". As far as I can tell in this case, NYT's robots.txt of the time was obeyed, which permitted automated processing of all but one specific article for some reason.
Probably. The question for the courts to decide, then, is how much use is considered fair use.
Isn't this what Mistral AI did?
Not sure about the rest of the world, but at least for US content I don't think any company would publish that LLM.
That's like 40 years before the civil rights movement, and right about the time of the Tulsa massacre.
It's right around when women got the right to vote.
Trying to get it to not say anything horrible under modern standards seems fraught with issues. I don't know if it would even understand something like "don't be racist", given the context it was trained on.
"I say to you that the VCR is to the American film producer and the American public as the Boston strangler is to the woman home alone."
Then we invented from whole cloth reasons why they were perfectly OK because there was a ton of money to be made and everyone would actually be better off if the VCR was a thing and everyone knew it because it ended up argued after millions of VCRs were already in households.
The LLMs built and trained 10 years from now will be much more advanced. Will it be possible for anyone to prove if specific content was used/not used in the training? If the courts rule against fair use, it will be a minefield to enforce.
Most copyrighted works will specifically say that the customer / user is prohibited from storing and reproducing those works.
It's a complicated area, not clear cut at all
In this case, then, regurgitation is the problem then, not the fact that it was ‘read’.
If the models ensured that probability of regurgitation is near-zero, would that be ok?
Perhaps you personally would be fine with it. But would it be ok for a court declaring that someone has no recourse, and must accept such an uncompensated risk?
I feel like I'm going to be saying a lot in the coming years, as more and more people's brains get broken by false anthropomorphization.
More here: https://en.wikipedia.org/wiki/Nitter
Interestingly, I think this going through would actually help the big players (Google/MS/Apple) a lot in the medium term. They sit on huge amounts of training data from their own services and also have the money to acquire copyrighted material, while everyone else would have to build on datasets under a permissive license (like Wikipedia), since scraped data would be a minefield.
I would then offer to pay damages of 3x revenue for all historic requests that resulted in 50+ words of Times articles to be reproduced.
And an analysis of the logs will probably show that only happened tens of times and total revenue from those requests was like $3.
In interactions with AI systems like ChatGPT, the user's intent and query nature significantly influence the AI's output. Regardless of user understanding, the fact that they engage with a generative prediction model is key. This novel use case is a first of its kind.
Similar to legal searches, where responsibility rests with the searcher, the intent behind AI queries — whether for research, information, or to demonstrate something about AI functionality — dictates the AI's response. The distinction between the AI's training data (e.g., articles from The Times) and the model's outputs is critical for evaluating transformative use. Each instance must be considered individually to ascertain this.
Just as intent is pivotal in legal contexts, it's also relevant when users interact with AI. If a user seeks verbatim content from sources like The NYT, responsibility could shift more towards them. This raises ethical questions about the intent behind the NYT's use of OpenAI's services.
If the NYT utilized OpenAI's services contrary to terms of use, such as for illegal activities or spreading misleading information, it would constitute a violation. Similarly, manipulating outputs to damage OpenAI's reputation would also breach these terms.
In copyright law, responsibility typically lies with the entity making copies or distributions. However, AI complicates this, as it generates content from various inputs. The user essentially initiates and guides this process through their prompts.
OpenAI could respond by filtering any NYT content and requiring users to agree to a legally binding contract. This contract would stipulate conditions for content usage in line with fair use principles, emphasizing a joint responsibility between the user and OpenAI. Such an approach aligns with other services where access is contingent on agreeing to specific terms.
gee why didn't the piratebay think of that
"we're just a search engine, if you search for Disney movies that's your fault, you even agreed to it in our terms of use!"
ChatGPT wouldn't be useful at all without mass unauthorised use of copyrighted material (same as the piratebay)
just ChatGPT has been tarted up with a facade of respectability
Things are about to get worse for generative AI - https://news.ycombinator.com/item?id=38814093 - Dec 2023 (548 comments)
Also:
NY Times copyright suit wants OpenAI to delete all GPT instances - https://news.ycombinator.com/item?id=38790255 - Dec 2023 (870 comments)
NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 - Dec 2023 (84 comments)
The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 - Dec 2023 (861 comments)
The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/item?id=38781863 - Dec 2023 (11 comments)
Plenty of discussion days ago about this and the case:
https://news.ycombinator.com/item?id=38781941
https://techcrunch.com/2023/04/28/openai-funding-valuation-c...
https://techcrunch.com/2022/03/02/snowflake-acquires-streaml...
Call Microsoft.
I bet the NYT relies on MS's OS, Software and Cloud services.
That's a lot more leverage.
Microsoft is the first defendant listed. You can read the Times' claims against Microsoft.
So that’s a window into this guy’s personality.
Just your usual smiling trendy unkempt beard sociopath corporate marketer in post-corporate business casual trying to hide as a folksy warrior of the people.
Or maybe he's an AI generated fuckface profile that cranks out ai generated propaganda for the AI corporation. Well, if he's real he should be prepared to be replaced by one soon, because word salad bullshittery is incontrovertibly the BEST thing llms are capable of.
Many individual artists have similar claims [0].
Now NYT has.
Individuals cannot foot the bill for lawyers so they have no choice but to see it all while being unable to do anything. NYT apparently has the muscle to pull it off.
Morality or ethics don't stand a chance here, it is all about who can march the largest army of lawyers which clearly, is the Tech bro clan.
I'm sad for all the creators. And of course, those marketing it as "synthetic intelligence" are on a whole another level.
There’s diminishing returns in this regard. Me suing Google has this dynamic, but the NYT is large enough to hire perfectly effective legal counsel. 1000 top lawyers won’t necessarily beat 100 top lawyers.
And please spare me the humans learn too nonsense. Computers are not humans.
The NYT is asserting a drastic expansion of copyright, to cover not just the specific expression of ideas but the ideas themselves.
We should be wary of AI companies, and copyright is woefully unprepared for our new world, but the answer is not to change the system so creators retain infinite control over who can even remember their work.
Show me a human can remember even half of what GPT4 was trained on and I’ll concede the point.
Those all influenced you. You may not have perfect memory, but neither does GPT4. You may rememeber a tiny subset of that data perfectly, just like GPT4.
The inanity here is in seeking to extend copyright beyond reproduction and into learning and cultural accretion. It’s crazy to me that anyone would look at the past 100 years of copyright and think that what we need publishers with legal control over reading.
Machine learning, including foundation models and web-scale pretraining, has widespread uncontroversially beneficial applications (defect detection, language translation, spam/DDoS filtering, agriculture/weather/logistics modelling, etc.). The federal government has invested billions in AI, and is desperately trying to prevent China from taking the lead. NYT aren't really in a position here to dig their heels in with "not our problem if our deletion demand makes training AI in the US infeasible" - they likely know that's not going to fly, and are instead using it to prompt negotiation for licensing their content.
Let’s imagine a really obvious case of copyright infringement - where a shady publisher buys up copywritten books, assembles them into a private library, then prints and sells copies. The infringement is the “prints and sells”, right? Not the “assembles a private library”?
If we accept that ChatGPT is committing copyright infringement (which at least seems true if it returns big chunks of the copywritten material verbatim), I don’t think you can get from there to compelling deletion of the model.
LLMs are basically an optimization of infinite monkeys using statistics.
I can see how outputs can be infringing copyright, but I’m having a very hard time seeing how the weights can be. It feels like saying a musician is infringing copyright if they understand a song well enough to be able to create an infringing reproduction, even if they don’t.
Preach it louder! I think most of them (certainly Huggingface and Mistral) will be ready to help on this. Let's hope that big tech companies are savvy at getting better lawyers who can more clearly articulate the pro-LLM position to lawmakers.
You may have a point for verbatim regurgitation of code, but I’m having a hard time seeing a violation in ingestion code and learning how it works.
BSD, MIT, several others. Without a license you cannot read the code.
there have at times been regimes that legally restricted the mere reading of published works, but such policies have generally been considered repugnant to liberal democracy and indeed liberalism in general
It's gonna take Robert Kardashian to get that through
I suppose they could file an amicus brief restating what openai's lawyers will say, but beyond that what are they supposed to do?
Before someone comes up with the argument that LLMs are creative and genuine, yes as a technique they are. The art form, the theory and science is marvellous result of human ingenuity. Hats off.
But the end result, is not.
Just like tape recorder is a great invention but the pirated music recorded on it is not.