Does the Common Crawl data already take care of the copyright issue? Or else how does the LAION crawler deal with that problem?
I mean, it's not too hard to write an image crawler. Also, scaling it up it a bit of a challenge, but it's a technical one. But the real difficulty is how to deal with all the legal strings attached...