This shows pretty clearly that the models do retain and return large chunks of texts exactly how they read them.
This shows pretty clearly that the models do retain and return large chunks of texts exactly how they read them.
One model is trained on copyrighted works in a jurisdiction where this is allowed and outputs "transformative" summaries of book chapters. This serves as training data for the deployed model.
A cover band who plays Beatles songs = great An artist who paints you a picture in the style of so-and-so = great
An AI who is trained on Beatles songs and can write new ones = exploitative, stealing, etc. An AI who paints you a picture in the style of so-and-so = get the pitchforks, Big Tech wants to kill art!
Has to pay the Beatles for the pleasure of doing so.
Sure, some copyrighted works ended up in the Pile by accident. You can download these directly, without the elaborate "poem" trick.