Google scanned many books quite a while ago, probably way more than LibGen. Are they good to use them for training?
What I'm wondering is if they, or others, have trained models on pirated content that has flowed through their networks?
I’m surprised Google hasn’t hit its competitors harder with the fact that they actually got permission to scan books from its partner libraries and Facebook and OpenAI just torrented books2/books3, but I guess they have aligned incentive to benefit from a legal framework that doesn’t look to closely at how you went about collecting source material