It's a good thing they don't train on copyrighted material from jurisdictions that have something other than "fair use", such as Australia's "fair dealing". Otherwise, they'd have to argue that their use of such copyrighted material doesn't "usurp either the market of the original work or a derivative market".
https://addisons.com/knowledge/insights/fair-use-or-fair-dea...
Do they? "Fair Use" is an affirmative defense, so the only time we're going to get into that is in a court case, where it'll be tested through legal means.
I would say it's even more nuanced: if LLM training involves merely reading a dataset, but it is not strictly necessary to copy, or even store it verbatim to be useful, then does it even fall under copyright protection at all? A lot of computer-based data processing is already immune to copyright issues; you can place a webpage into a server-based cache or CDN, you can stream it across a network, you can cache it in RAM or local storage, you can make backups of things, and all these processing uses don't fall afoul of copyright.
So I would say that we're going to watch the LLM trainers say that the models aren't storing copies at all, and that seems an even stronger defense than "Fair Use". It is a strange copyright protection indeed that explicitly or implicitly prohibits certain types of machine readings, while allowing many others.
It is a fair point. Companies contend what they always contend, which is their position in an argument; they do so forcefully and regardless of the reality on the ground. Companies are basically modeled after opportunists.
Copyright includes the creation of derivative works, not just literally copying the source material.
For instance, imagine I read a novel, then I decide to write my own, unauthorized sequel to it. It's not a literal "copy" of the original material - it's my own original text, but obviously a derivative work of the original material. Under copyright law, that would be infringement - I would be sued if I tried to sell that. (Yes, that means fanfiction is infringing, but most rights holders have wisely decided to look the other way on that, as long as it's non-commercial.)
This is what people who claim AI is infringing are worried about. Not that the AI has a literal copy of the source material in its training data, but that the training data can be used to produce a derivative work.
I could write a (crappy) fanfic of the Lord of the Rings without directly referencing the books/movies. And that doesn't mean I have a complete copy of the books/movies in my head - that isn't how memory works. Until now, creating a derivative work without directly using the source material was something only humans could do. This is completely uncharted legal territory.
As far as I know, licenses can discriminate on whatever constraint they want to[1].
A license that is basically "This work is for $FOO only. If you want a $BAR license please contact us." is perfectly legal right now!
There is no restriction in law that $FOO cannot be "human consumption" and $BAR cannot be ""machine consumption".
If the AI companies are not arguing the "fair use" argument, then they are arguing for AI companies having a special exemption for themselves carved out in law whereby a copyright holder is forced to license their works to the AI companies specifically.
IOW, their only recourse is to argue "fair use", because any other argument boils down to "please make a special exemption for us in copyright law, defying hundreds of years of precedent", which is a particularly hard sell.
[1]Excluding discriminating against protected classes.