The more disregard a company has for intellectual property rights, the more data they can use.
Google had far more to lose from a "copyright? lol" approach than OpenAI did.
Google had far more to lose from a "copyright? lol" approach than OpenAI did.
The key questions are around "fair use". Part of the US doctrine of fair use is "the effect of the use upon the potential market for or value of the copyrighted work" - so one big question here is whether a model has a negative impact on the market for the copyrighted work it was trained on.
Take a look at https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... - bullet points 2 and 4 on pages 2/3 are about training data. Bullet point 5 is the Bing RAG thing.
The company that scrapes trillions of web pages has an issue with copyright?