I would speculate this is true of all the leading commercial LLM models. Don’t have enough training data? Just steal some!
During the early Llama 1 days The Pile dataset was in heavy use by many. Bit later people figured out that a subset of it - Books 3 - was especially problematic.
I’m guessing all the big houses threw that piece out in later models since it’s extra radioactive
- if you are an individual then it's called "pirating copyrighted work"
- if you are a multi-billion dollar corporation then it's called "use of uncleared material for training"
Zuckerberg murdered some old library books to train a model. Zuckerberg genocided training data!
Heck, everyone who read your comment here stole it. I'm so sorry for your loss.
This is just that with lots of levels of indirection.