OpenAI's Motion to Dismiss Copyright Claims Rejected by Judge
arstechnica.com
arstechnica.com
Also, it’s easier to remove copyright material if it’s not all crammed into an LLM first. Eg. If someone wanted to remove their website from Google, you can do that incrementally without rebuilding the entire index, whereas it’s a lot harder post-LLM (post-processing is probabilistic at best).
Websites tend to be okay with it because they accrue a benefit of Google’s crawling - they get traffic back. When websites don’t feel that Google keeps the traffic for themselves, websites tend to get upset https://www.theregister.com/2020/03/11/yelp_congress_google/
LLM training just takes and keeps all benefit to themselves. Wikipedia (or news site) get no traffic or anything back in return.