NYTimes, CNN and ABC block OpenAI’s GPTBot web crawler from accessing content
theguardian.com
theguardian.com
Sites scramble to block ChatGPT web crawler after instructions emerge https://news.ycombinator.com/item?id=37094463
GPTBot – OpenAI’s Web Crawler https://news.ycombinator.com/item?id=37030568
Of course OpenAI probably still has a full dump of LexisNexis or something.
People assume you need ingest a bunch of an author's writing to have the LLM produce writing in that style... I think people seriously underestimate the level of inference these LLMs are capable off.
https://chat.openai.com/share/58b38f84-4ddd-4147-989a-c8cbf2...
For example:
> Inference: A scholar specializing in narratology might observe that the article employs a classical Aristotelian narrative arc, replete with tension ("nail-biting 45 minutes") and resolution ("heartwarming testament"). This is more typical of storytelling than journalistic reporting, which often employs an "inverted pyramid" structure to relay the most crucial information first.
To write an analysis like that, the LLM isn't regurgitating tokens in a way that matches the NYT, it's close to regurgitating academic analyses of writing. You could train the model without a single NYT article and it'd still be able to produce that guidance and then apply it to novel writing pieces to get something NYT-like.
Also, the lesson of Google News was that news is a commodity. The world is full of websites reporting on the same events and giving those reports away for free. They aren't all going to block GPTBot.
And then finally this might even be a good thing for OpenAI. One of the criticisms of ChatGPT is that it has a left wing bias, probably a result of training on left wing biased news media. If NYTimes/CNN/ABC stuff no longer shows up in the training set, ChatGPT's views may come more into line with that of the average person.
https://platform.openai.com/docs/gptbot/disallowing-gptbot
I would not be the slightest bit surprised if some "error" prevented this from working as intended, but hey, what do I know?
The scrapers are all rewriting their privacy policies to just say "all your text is belong to us" so I don't see much recourse for the media outlets, aside from being large with lots of aggressive IP attorneys.
And what becomes of all the archive.is pages we're feeding to them anyway?