Of course OpenAI probably still has a full dump of LexisNexis or something.
People assume you need ingest a bunch of an author's writing to have the LLM produce writing in that style... I think people seriously underestimate the level of inference these LLMs are capable off.
https://chat.openai.com/share/58b38f84-4ddd-4147-989a-c8cbf2...
For example:
> Inference: A scholar specializing in narratology might observe that the article employs a classical Aristotelian narrative arc, replete with tension ("nail-biting 45 minutes") and resolution ("heartwarming testament"). This is more typical of storytelling than journalistic reporting, which often employs an "inverted pyramid" structure to relay the most crucial information first.
To write an analysis like that, the LLM isn't regurgitating tokens in a way that matches the NYT, it's close to regurgitating academic analyses of writing. You could train the model without a single NYT article and it'd still be able to produce that guidance and then apply it to novel writing pieces to get something NYT-like.
Also, the lesson of Google News was that news is a commodity. The world is full of websites reporting on the same events and giving those reports away for free. They aren't all going to block GPTBot.
And then finally this might even be a good thing for OpenAI. One of the criticisms of ChatGPT is that it has a left wing bias, probably a result of training on left wing biased news media. If NYTimes/CNN/ABC stuff no longer shows up in the training set, ChatGPT's views may come more into line with that of the average person.