Notwithstanding the provisions of sections 17 U.S.C. § 106 and 17 U.S.C. § 106A, the fair use of a copyrighted work, including such use by reproduction in copies or phonorecords or by any other means specified by that section, for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include:
1. the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
2. the nature of the copyrighted work;
3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
4. the effect of the use upon the potential market for or value of the copyrighted work.
The fact that a work is unpublished shall not itself bar a finding of fair use if such finding is made upon consideration of all the above factors---
Let's polish over the "scholarship, or research" clause for a moment, because none of the commercial entities are doing it for research. Assuming the fair use clause didn't fail there, let's focus on two important points:
3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
4. the effect of the use upon the potential market for or value of the copyrighted work.
First of all, an LLM remembers almost all of its training material verbatim, or almost verbatim, which is undoubtedly a substantial amount of the work ingested*, and secondly, the effect of such systems on the whole market or job ecosystems is disruptive. You can create something in the style of someone from your couch for free or $10/mo.Whatever you do, you can create derivative works of the ingested works, or derivative works containing ingested parts verbatim at great speed with this "Shiny LLM thing".
A human neither can ingest that amount of information, nor retain all of it that perfectly, hence limiting the inspiration from these works, and forces one to add their own original ideas or inspirations rooted in other sensory inputs (feelings, sight, own experiences, etc.), not reducing the value of the work used, or not damaging the creator of the work the creator got inspiration.
Incidentally, we have a word for cases this inspiration gets a bit too far: "plagiarism".
From that point, an LLM, which remembers everything verbatim and mixing them together is arguably making plagiarism from many sources to create their own works, while being unable to cite them accurately. A very useful property, indeed.
*: I know how an LLM stores its data in its weights, but even without the data itself, the weights can construct the training data verbatim in many cases, so it stores its training data. Not in its original form, but in a derivative and reversible form, not unlike compression.
[0]: https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors