I’ve tried to dig a bit through the libgen fiction archive. Surprising to me, I found that the vast majority was Romance novels. I’ve since heard this from others too that most published books are Romance.
If this collection is trying to solve the problem of having large tagged and sorted data for training, even if it can’t be commercialized, it might also be introducing another where the data still needs to be weighted and filtered. Not to knock on Romance specifically but even training on all scientific papers ever written, whether peer-reviewed, highly cited, retracted, debunked, etc. might explain some hallucinations it ends up making.