I remember a few years ago Google had a big blog post about how they’d injected some fake search results to catch Bing scraping like this. It certainly seemed like they thought it was a bad thing back then.
I remember a few years ago Google had a big blog post about how they’d injected some fake search results to catch Bing scraping like this. It certainly seemed like they thought it was a bad thing back then.
Google noticed this, and for unique/low-traffic search terms, was able to synthetically generate enough "fake" traffic that the "fake" traffic became the dominating signal for those terms, and therefore Bing started directing users to the fake results.
The bad thing here IMO is the level of tracking of users via this toolbar, but fundamentally this seems as bad as any other digital fingerprinting or advertiser tracking as anything else that's become common on the web. This is not to excuse Microsoft for doing a bad thing, but it really had very little to do with "scraping Google", which somehow became the popular media takeaway for this.
EDIT: a decent contemporaneous article in Wired: https://www.wired.com/2011/02/bing-copies-google/, mentioning how Microsoft was using the clickstream data from the browser/toolbar, not scraping Google results per se
I wonder how common misunderstandings like this can be prevented or treated.
If you specify that you don't want Google to take snippets, but a third party ignores the robots.txt or meta tags and has a permissive one of its own Google can just scrape it from there.
A missing detail here is what did the textbox for the fake celebrities' net worth on Google link to?
* https://www.google.com/search?q=larry+david+net+worth -> featured snippet from Wikipedia, showing an estimate
* https://www.google.com/search?q=bill+gates+net+worth -> database style response, showing answer with no link
https://techcrunch.com/2020/08/11/court-dismisses-genius-law...