> It's not clear. I do not consider training an LLM on publicly disseminated text to be "morally on the level of theft." Stealing my car, or even a pen off my desk, is a much more reprehensible action than slurping everything I've shared in public into an LLM training process, purely IMHO.
The theft is that of effort, in the exact same (or a worse) sense as pirating media or stealing IP from a company.
It takes effort to write. That effort is being stolen by an LLM during the training process - the LLM cannot possibly exist without the work done by the authors who wrote content that it is being trained by, and the LLM can also be used to automate away those authors' ability to do work (and jobs) by replacing them. Which is worse - to have your car stolen (which is very bad, I'm not arguing that it isn't), or to lose your job, and not being able to afford anything?
Alternatively, if you believe that it's not bad to take someone's effort without their consent and without compensating them for it, then you shouldn't object to your employer withholding wages from you, or a client refusing to pay you, on the same principle.
> People or corporations (which are usually treated as person-like) operate training processes and are morally and legally responsible for them. I believe training an LLM is "a human/corporation doing something" with a piece of content.
That's not reasonable, and most people do not share your opinion (including the relevant group, which is the authors of the content being trained on). That's equivalent to saying that a human writing a program to perfectly reproduce a copyrighted work (e.g. print out the complete text of Harry Potter) is a human "doing something" with that copyrighted work (in the same class as reading Harry Potter).
> Whether I let an AI agent read your blog post or whether I write a program to read it into an LLM fine tuning process seems immaterial to me.
Those are categorically different. The vast majority of the population (again, including those writing the works that are being trained on without their consent) will agree that they are categorically different and incomparable, and they are logically, legally, and morally distinct.
> One of those exceptions (in many jurisdictions) is making temporary copies of data to use in a computational process. For example, browser caching, buffering, or transient storage during compression/decompression.
To use in specific computational processes for which you do not store the output because the output is subject to the same copyright laws. The implicit premise when you talk about training is that you're going to save the trained model, so this obviously doesn't apply, in the same sense that if you take a copyrighted work and transcode it, the transcoded output is subject to the exact same set of copyright laws as the original.
> not those merely training LLMs which is not, in and of itself, violating any laws I can establish.
That's the "law is morality" fallacy. Morally, this is clearly wrong, the point of the copyright system is to prevent exactly things like this from happening. The courts have not yet decided whether training an LLM is "copying" a copyrighted work, but if they do, then it's clearly illegal.