"If I feed all 7 books of Harry Potter into an AI, and only that, and then it spits out something - That's clearly a derivative work." does not match the current legal interpretation in the industry.
The general assumption is that machine learning models trained on some data are (usually - details matter) not considered derivative work from that data; so unless you have specific contractual restrictions (e.g. you got the date bacause you have a contract with the data owner saying what you'll do/not do with the models derived from it) then copyright does not prohibit you from using and distributing these models. That's the case even if it's trained on a particular dataset with a single, clear copyright owner.
It probably starts with old precedent on copyrightability of statistics derived from creative work - things like word frequency, most common words, etc are not considered derivative works.
Then we have statistical language models, of the kind that were used in statistical machine translation before the neural approaches overtook everything - and again, the established interpretation there is that statical models trained on a corpus are not considered derivative work from that data, because they essentially are a trivial extension of word sequence frequency counts... but they already are in the category of "feed all 7 books of Harry Potter into an AI, and only that, and then it spits out something", e.g. a simple hidden Markov model or trigram language model from Harry Potter books will easily hallucinate (very lousy quality) Potter-themed text, but the model that can do that is not considered derivative work.
Though this all may vary in different jurisdictions.