The copyright situation around all this is very... interesting. Pretty clear that this dataset is not legal but what about resulting models? What if the texts actually where bought 'properly'?
Edit: also, I don’t believe court decisions can be enforced retroactively so existing LLMs would be safe but I’m most definitely not a lawyer.