Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "enshittification". (The OALD 7ed in it's foreword even states "Corpus analysis now makes it possible to draw authentic examples from a vast range of attested contemporary usage. A concordance will display hundreds or thousands of them to choose from.")
I suspect that the inclusion of a few sentences -- especially those that introduce a new word or usage of a word -- are fair use, but the inclusion of the entire texts is not.
This then brings up an interesting point where the computer scientists/linguists developing tools like WordNet or other NLP databases would be at an advantage to those that take the approach of throwing a lot of data into a neural network and hoping for the best. Yes, it is a lot more work/effort to develop those NLP databases, but in the end they may end up being more robust, especially around the question of copyright.