The reason is the staff at the NYT appear to be very well versed in the technical tricks people use to gain access.
The reason is the staff at the NYT appear to be very well versed in the technical tricks people use to gain access.
Synchronously verifying it, would probably be too slow.
You can verify googlebot authenticity by doing a reverse dns lookup, then checking that reverse dns name resolves correctly to the expected IP address[0].
[0]: https://developers.google.com/search/docs/crawling-indexing/...
Why would it be slow? There is a JSON documenbt that lists all IP ranges on the same page you linked to:
https://developers.google.com/static/search/apis/ipranges/go...
But if I were implementing filtering, I might prefer a solution that doesn't require keeping a whitelist up to date.
Not to mention the possibility of just filling up the banned IP table.
With recommender systems, attention graph modeling, etc. it'd probably be a perfect information ingestion and curation engine. And nobody else could modify the algorithm on my behalf.
I don’t know about paraphrased versions but it would need to handle content revisions by the publisher somehow.
It appears anyone can read any new NYT article in the Internet Archive. I use a text-only browser. I am not using Javascript. I do not send a User-Agent header. Don't take my word for it. Here is an example:
https://web.archive.org/web/20240603124838if_/https://www.ny...
If I am not mistaken the NYT recently had their entire private Github repository, a very large one, made public without their consent. This despite the staff at the NYT being "well-versed" in whatever it is the HN commenter thinks they are well versed in.
Because I had to learn more, sounds like a pretty bad breach. But I’m still pretty impressed by NYTs technical staff for the most part for the things they do accomplish, like the interactive web design of some very complicated data visualizations.
You might be thinking of Mike Bostock, the creator of D3.js and ObservableHQ.com, who led data visualization work at NYT for several years [1]. I'm not sure if they have people of that magnitude working for them now.
Which browser is this?
Thanks for the tip!
Edit: Just read your comment again. I assume that’s exactly what you meant.
Example for LA libraries: https://www.lapl.org/new-york-times-digital
Your local library might have a similar offering specifically for the NYT.