Actually, I think they could comply by providing a list of every author in the dataset, though this is following the letter rather than the spirit.
They could also do research on ways to get the model to return the top ten influential works for some output, and make a legal argument that this is a best effort given technical challenges with tracing every source.
For the first point, I get that by reading the license at this image:
https://commons.m.wikimedia.org/wiki/File:3H8A7368.jpg
“attribution – You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
share alike – If you remix, transform, or build upon the material, you must distribute your contributions under the same or compatible license as the original.”
So, output has to have the same license and worst case the image is accompanied by a link to a list of every author in the dataset. This is a far cry from the “this research would be useless if copyright was enforced” as some people suggest.
Here’s an open dataset I found which does not require attribution:
https://www.pexels.com/creative-commons-images/
And Wikimedia commons, some of which require attribution:
https://commons.m.wikimedia.org/wiki/Category:Images
And it is easy to go take a 4k video camera and start collecting tens of thousands of frames of your own images.
My point is that people are throwing up their hands saying oh, respecting the copyright of artists is impossible. But it feels very unfair that these huge companies are walking all over the copyright of small artists, but if we took their code to re use their lawyers would sink us overnight. This upsets a lot of people and it’s a bad look. I don’t actually like copyright but if everyone else has to follow the rules I don’t like giving them a free pass.