Plus other sites that link to the content could also give away it's date of creation, which is out of the control of the AI content.
I believe I learned about it through HN, and it was this blog post: https://hallofdreams.org/posts/physicsforums/
It kind of reminds me of why some people really covet older accounts when they are trying to do a social engineering attack.
According to the article, it was the founder himself who was doing this.
None of these documents were actually published on the web by then, incl., a Watergate PDF bearing date of Nov 21, 1974 - almost 20 years before PDF format got released. Of course, WWW itself started in 1991.
Google Search's date filter is useful for finding documents about historical topics, but unreliable for proving when information actually became publicly available online.
https://www.google.com/search?q=site%3Achatgpt.com&tbs=cdr%3...
So it looks like Google uses inferred dates over its own indexing timestamps, even for recently crawled pages from domains that didn't exist during the claimed date range.
I wonder why they do that when they could use time of first indexing instead.