No need for surprises! It is publicly known that the corpus of 'shadow libraries' such as Library Genesis and Anna's Archive were specifically and manually requested by at least NVIDIA for their training data [1], used by Google in their training [2], downloaded by Meta employees [3] etc.
[1] https://news.ycombinator.com/item?id=46572846
[2] https://www.theguardian.com/technology/2023/apr/20/fresh-con...
[3] https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o...
"Researchers Extract Nearly Entire Harry Potter Book From Commercial LLMs"
https://www.aitechsuite.com/ai-news/ai-shock-researchers-ext...
So far, courts are siding with the "fair use" argument. No need to exclude any data.
https://natlawreview.com/article/anthropic-and-meta-fair-use...
"Even if LLM training is fair use, AI companies face potential liability for unauthorized copying and distribution. The extent of that liability and any damages remain unresolved."
https://www.whitecase.com/insight-alert/two-california-distr...
This definitely raises an interesting question. It seems like a good chunk of popular literature (especially from the 2000s) exists online in big HTML files. Immediately to mind was House of Leaves, Infinite Jest, Harry Potter, basically any Stephen King book - they've all been posted at some point.
Do LLMS have a good way of inferring where knowledge from the context begins and knowledge from the training data ends?
Anna's Archive alone claims to currently publicly host 61,654,285 books, more than 1PB in total.
https://www.washingtonpost.com/technology/2026/01/27/anthrop...
Anthropic, specifically, ingested libraries of books by scanning and then disposing of them.
The plot of Good Will Hunting would like a word.