This analogy isn’t quite right, it’s more like if you trained a font generation AI using commercially licensed fonts, or trained a literature generating model on samples of copyrighted fiction.
The part that matters is that the model is being trained on the copyrighted features of the input, not the parts the copyright holder doesn’t care about.