You don't even need that. Even "just" OpenAI has a valuation large enough to just
buy several major publishers and data brokers to secure access to data if they need to license.
And while I expect NYT imagines that their archive is really valuable for training, they're just not that special in the sense that while they may have broken more stories on average than many others, and have had influential op eds etc., the ones that matters will have been cited and referenced and written about elsewhere - the irony is that by virtue of being so well known, their historically most important content is also less unique in terms of the accessibility of the information in it.
So while I'm sure OpenAI would love their archives, I'm also sure that if OpenAI and others have to license content and NYT end up being "difficult", OpenAI will just license content from (or buy) a suitably diverse portfolio of other papers instead.
In other words, beyond producing outright synthetic data, if AI companies are prevented from training on data they don't have a license to, the net effect will just be a scramble to buy licenses and/or buy companies that can provide sources of content, and the price for that content will be a lot lower than some of the people pursuing these copyright claims imagine.
In the end, if we go that route, all we'll have achieved as a society is creating massive moats protecting the companies already big enough to buy access to a broad enough set of content and made open models harder.