https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp....
However, it seems that there is a later case in the 2nd circuit:
https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol....
https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp....
However, it seems that there is a later case in the 2nd circuit:
https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol....
Peak was a repair business. MAI built computers (as in assembled/integrated; I think they were PCs) and had packaged an OS and some software presumably written or modified in-house along with the computer. MAI serviced the whole thing as a unit. So did Peak. MAI sued Peak for copyright infringement because Peak was taking computer repair/maintenance business away from MAI, under the theory that Peak employees operating their clients' MAI computers and software was copyright infringement. (There were other allegations of Peak having unlicensed copies of MAI's software internally, but that's not central to the lawsuit.)
If you have a piece of IP to use to train an IP model with, and you have legal right of access to use that piece of IP (for private purposes), MAI v. Peak doesn't cleanly apply.
MAI v. Peak is also 9th circuit only, and even without the poor reasoning, it should automatically be in doubt because the 9th circuit is notoriously friendly to IP interests, given that it covers Los Angeles.
I was only pointing out that the law is of the opinion that a copy is a copy is a copy, regardless of where it's made, or how long it exists for.
Other decisions come into play to save us, like Authors Guild v Google, where they said search engines could make copies, bringing Fair Use into the picture.
Personally, I think that creating the model is Fair Use, but anything produced by the model would need to be checked for a violation. I would treat it the same as if I went to Google Book Search, and copied the snippet it returned into my new book.
The license associated with the training data then becomes insanely important. Having the model reference back to the source data is even more important.
For example, training data with a CC BY license would be very different to CC BY-SA and CC BY-ND, and they all require the work produced by the model to have credit back to the original source to be publishable.
When an artist displays their work on DeviantArt or Artstation or whatever, they are allowing the general public to load it into memory. It's part of the license agreement they sign when they sign up for these services.
Fair Use applies to instances that would otherwise be copyright violations, i.e. unauthorized distribution.
When you sign up for a social media site you EXPLICITLY grant the site the rights to distribute it. You have expressly permitted it. It's a big difference!
There hasn't been an explicit decision for ML training, but everyone's assuming that Authors Guild v Google applies.
https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.....
CDNs operate under the control of the copyright owner, so they would be authorized.
Web browser caches are under the control of the recipient who has authorization to make a copy.