It has been frustrating.
So training must consider licencing where copyright material is used and not consume all data.
Your brain is not a model. You can not reproduce most of what you see. You're not "training" your brain by glancing at an image as your recall concerning that image will be terrible.
I'm sure the millions of people who violate copyright law daily with absolutely no repercussions care very much about that.
You cant setup a cinema and charge ticket for the movies you stole.
Its the money making side that matters - not individuals ij a private house
Am I violating copyright law because I am merely capable of producing a copy of something? Obviously not. Why should the model be?
> The current (legal) answer is "unclear".
European Union was ahead of times for once. The 2019 copyright directive, article 4, makes it legal to scrape the web and make and keep local copies of copyrighted works, for data mining purposes. Unless the copyright holders set up a machine readable exception (such as robots.txt file).
So legal in EU, "unclear" in US.
'copyright fair use' : https://copyrightalliance.org/faqs/what-is-fair-use/