There is something novel here.
Google Books created a huge online index of books, OCRing, compressing them, and transforming them. That was copyright infringement.
Just because I download a bunch of copyrighted files and run `tar c | gzip` over them does not mean I have new copyright.
Just because I download an image and convert it from png to jpg at 50% quality, throwing away about half the data, does not mean I have created new copyright.
AI models are giant lossy compression algorithms. They take text, tokenize it, and turn it into weights, and then inference is a weird form of decompression. See https://bellard.org/ts_zip/ for a logical extension to this.
I think this is the reason that the claim of LLM models being unencumbered by copyright is novel. Until now, a human had to do some creative transformation to transform a work, it could not simply be a computer algorithm that changed the format or compressed the input.