Hear, hear. Unfortunately this is more or less impossible given current copyright law.
Suppose you scrape libgen and turn it into training data, then you release the training data publicly. Since the vast majority of every book appears verbatim in the training data, is this sufficiently transformative?
I think yes, it is, because nobody is going to read those books from the training data. When I made books3, I felt it was important to render each book into high quality text. But it turns out that when you convert Jurassic Park into a text file, there's no good way to read it anymore. Good luck trying to bookmark wherever you left off -- it's all one gigantic file.
But nobody seems to agree. The Danish Rights Alliance (https://rettighedsalliancen.com/) aggressively DMCA'ed anyone that hosted books3, even going so far as to DMCA The Pile from academictorrents: https://academictorrents.com/details/0d366035664fdf51cfbe9f7... with the justification that ~100 copyrighted books appear in the training data, so therefore they have the right to DMCA. Right now most of the world seems to agree, but I'm hoping that opinion will shift as the years tick by. Surely no one can believe that a plain text document poses a serious threat of economic harm to the original author. So the question is whether the original author should be allowed to deny everyone else the right to transform their work into a form that machines can read.
For my part, I've been planning a books4 dataset, but this time similar to LAION: it's a script that spiders libgen torrents (https://libgen.gs/torrents/libgen/) and converts all the epubs into text files. That way, if LAION isn't infringing, then books4 can't be infringing either. (Of course, hosting the actual training data anywhere is pretty hard nowadays, but it should only take a few days to convert 38TB of libgen into ~2TB of plain text.)
This is the only way to create an open source competitor to ChatGPT.