the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?
the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?
you would have to ask a judge about that because its a novel legal question. obviously nobody anticipated this technology at the time it was written so ultimately it will have to be a court that decides how to apply existing laws.
Again there is no "compression of information" in deep learning.
i would strongly disagree. when you are training a model you are taking the information from a document and extracting the relationships between tokens and storing that information conglomerated with the same information from a massive amount of other documents. the model that results is a compressed form of all of the information from all of the documents where you have extracted and stored a synthesis of the relationships between the tokens in all of them. this is a lossy compression, but it does reproduce exact sequences of source documents in some cases, so the original information is stored there.
you can very plausibly argue that an LLM model trained on copyrighted material violates the copyright on every single copyrighted document that was fed to it.