NYT vs. GPT: Microsoft finds itself on the other side of an industry
geekwire.com
geekwire.com
At least with software, there are explicit licenses governing how you're presumably allowed to use them. Is the New York Times going to claim people are agreeing to a license before reading their news?
They don't need to. They have implicit, automatic copyright on the published original works. The lawsuit shows examples of GPT using the NYT work verbatim.
OpenAI does not distribute it models, so the violation of copywrite if found to be a violation will be limited to the amount of people who queried the LLM in the same way that NYT did.
How that copyright shakes out in the end (what your statements focus on) is yet to be determined.
I'm saying that between "you're violating copyright law" and "you're violating copyright law and you supposedly agreed you a license that contractually prevented you from doing this", the second sounds like a stronger claim.
So existing copyright law was sufficient for written media, but didn't quite work for software. OTOH it seems that GPL-style copyleft licenses are being used more by writers and artists, so it's not like these issues exist in a vacuum. I expect copyleft will be more appealing to artists going forward given Midjourney et al.
I don't think they should have a right to the records of people who lived before simply bc they underpaid, over worked human journalists to exploit their creativity - largely to make the "Newspaper Barons" rich.
I don't think they ought to have implicit, automatic copyright on older stuff. How old is my only debate.
What good is old news really? Real old is good for college papers but news from 4 years ago... who cares?
"NYT bits of the corpus" is understating things - I imagine getting rid of NYT articles would significantly impact the overall quality of GPT-4's responses in many unexpected domains.
The NYT might be better than average, but it's not so special that you can't replace it.
The long-term effect if NYT wins this case will be that the "established" players and companies with deep pockets will get a massive moat protecting them against open models and new entrants - in the long run it might well be well worth whatever a loss will cost them.
The NYT could prolly win under current law - due to the interpretation of how GPT is reproducing the articles. The judge that may ultimately decide the future of how AI will fall for now in the US, will likely know very little about AI/LLMs.
I don't think reproducing a current article is that big deal tbh - it just retrieved the info from itself rather than the NYT bc it had already saved the info before and can't forget things. That's humanely possible by anyone with a photographic memory - it's more similar to recollection than infringement. The NYT is pissed the LLMs learned from their data - you can't stop something from learning - especially after the fact.
Obviously, the user intent not machine intent is key also. No infringement can happen without a prompt that suggests copyrighted material - easy to do or not, requires a person nonetheless.
The technical aspects of this is where they have a chance, taking a step back, it all just looks kind of silly.
You use the notes I gave you verbatim in your paper and are found guilty of plagiarism.
You obviously didn't realize the notes I gave you on the spot were from a textbook - that wasn't what you were looking for. I didn't realize you were just going to use whatever I wrote in your paper.
Did I commit plagiarism or did you?