I've done some work with corpus linguistics and quantitative linguistics, and large parts of these disciplines essentially are about facts derived from books in some manner. Modern approaches tend to involve machine learning, deep neural networks and other things fashionable on hackernews, but in general that's an old, traditional area that was working on facts derived from books for decades before the "ML era".
To work on facts derived from books, we're sourcing all kinds of books and other written language, such as newspapers. Some publishers and authors are cooperative and helpful for such research, some are uncooperative and prefer to intentionally make working on their sources difficult - but in any case, even in the case of disagreement and conflict there's no "IP war", the conflict in our case tends to be about practical convenience of access, not about IP, because they don't really have a leg to stand on in claiming a copyright violation. They hold the copyright on the original text, which gives them certain exclusive rights, there's a bunch of intermediary data that we can't make available to public without their permission, but these rights don't extend to facts derived from that text, and we legally don't need their permission to work on, analyze, transform, publish and use stuff based on facts in the text or facts about the text, we can do that openly even if they've explicitly made it clear that they don't want us to do that. That's nothing new, that's established law that probably predates modern computers.