I suspect ChatGPT is using a form of clean-room design to keep copyrighted material out of the training set of deployed models.
One model is trained on copyrighted works in a jurisdiction where this is allowed and outputs "transformative" summaries of book chapters. This serves as training data for the deployed model.