This way WSJ would be showing the same (www.wsj.com/Bezos-with-hair) to every user.
*based on the video
That's a new one. Buy a subscription to the WSJ now for a song.
Google's rules and help documents are spread all over, but here's some from Google News about subscriptions:
"If you prefer this option, please display a snippet of your article that is at least 80 words long and includes either an excerpt or a summary of the specific article. Since we do not permit "cloaking" -- the practice of showing Googlebot a full version of your article while showing users the subscription or registration version -- we will only crawl and display your content based on the article snippets you provide."
edit, forgot link: https://support.google.com/news/publisher/answer/40543?hl=en
In the video the guy explains that while you are not allowed to treat gbots request in a way no other people are treated, it is okay to differentiate between "boxes" of users. In their example it is the country USA, but if you define a "google user country" and let all users of google in it is ok to bundle the gbot with those. Grey area for sure, but might makes right.
"So geolocation, that is, looking at the IP address and reacting to that-- is totally, fine, as long as you're not reacting specifically to the IP of just Googlebot, just that very narrow range".
Also, they will crawl you from an unusual IP using a user-agent that doesn't say it's Google. And when that happens, and you deny access to undercover-Googlebot, but allow Googlebot in full uniform, you'll be penalized for cloaking.
They would have to react to just gbot when constructing the special google urls when the bot is crawling the site, tho.
This how they were originally handling it, before February. It would display if you visit the link from Google, or set your refer to it looks like it (this is why HN has the "web" link under articles), even if you weren't a subscriber. It's allowed, because regular users coming from Google do see the same thing as Googlebot.
WSJ have since changed that, so only subscribers can view articles, and you no longer get a "free click" coming from Google Search as Google calls it. They now show a short snippet, and are following guidelines to be labeled a “subscription” service by Google Search. This caused their rankings to drop below being a "free" news source though. But it's not nearly as bad as if they had cloaked Google.
For instance, if you happened to be a logged WSJ user, then google could show you the result based on your cookies.
I did not say that.
I said google could display the results to which you have access, based on your cookies. Assuming they had access to crawl the whole article.
It is not possible for Google to get access to cookies for other sites, anyway. This is a pretty fundamentally important security restriction that browsers implement to protect you from nefarious sites. So it isn't possible for Google to know which sites you have paid accounts with unless you explicitly tell it.
How would that work? Would Google create or be given a login for WSJ? That would benefit WSJ as an incumbent news provider at the expense of startups too new or small to get special treatment from Google.
How would a new website indicate to all search crawlers (so google doesn't befit at the expense of other search engines) how to get access to it's pay-for content in such a way that end-users can not also pretend to be seach crawlers and get access to the same content?
Sure, some people could set up GCP instances and proxy it, but that's a very tiny percentage of people.
DNS lookups are far more expensive to perform than an IP filter, and couldn't be done in realtime. So WSJ would have to set up a system where they regularly find all Googlebot referers in their logs that were rejected, do DNS lookups on the IPs, and add any that were valid to a whitelist so that they won't get rejected again. This will cause new Googlebot IPs to get rejected until the whitelist is updated, hurting indexing and ranking. The WSJ would also have to go through their whitelist regularly and do DNS lookups to verify that all of those IPs are still valid Googlebot IPs, and remove any that aren't valid anymore. That opens a window for invalid IPs to continue to get access, which may or may not be a problem depending on how often IPs change and where they get reassigned to.
The IP whitelist would need to be distributed to WSJ's webserver farms and used to update firewall rules, in an automated way that may or may not integrate with how that stuff is currently managed. (Generally, those rules would be tightly controlled in a big org like the WSJ.) The HTTP access log gathering from the farms and their analysis would also need to be automated, which again might be a management issue if the logs contain anything sensitive. (Like, I don't know, records of particular individuals reading particular stories which certain government agencies might be interested in acquiring without the hassle of a warrant.)
So yeah, there's a way to find out if an IP belongs to Googlebot. That's a long way from a manageable filtering solution at the WSJ's scale, even if Google wouldn't penalize them for doing it, which they would.