Problems caused by Google crawlers
hackerfactor.com
hackerfactor.com
see:
https://www.hackerfactor.com/blog/index.php?/archives/678-Bo...
https://www.hackerfactor.com/blog/index.php?/archives/762-At...
https://www.hackerfactor.com/blog/index.php?/archives/775-Sc...
https://www.hackerfactor.com/blog/index.php?/archives/777-St...
https://www.hackerfactor.com/blog/index.php?/archives/779-Be...
I've done lots of portscanning of the entire internet, this always generates vast amounts of abuse emails. Many of these these emails are xenophobic rants from obviously distressed people staring at their access.logs, thinking they're being targeted by the Russian government or something.
I can't help but wonder if this might be a similar case.
The fact that he says something like "a Google administrator tracked down a rogue employee and made him stop" makes it clear he's paranoid (or malicious).
Or sensationalism. (it did land a top spot on HN)
[0] https://www.hackerfactor.com/blog/index.php?/archives/484-Go...
I've never seen anyone who doesn't fit this stereotype posting analysis of URL hits. It's kind of seems that only this kind of person (which I actually like a lot — these guys are very fun when you have common interests) would be so anal about this.
This is just normal background noise on the internet, one has to be pretty paranoid to assign any significance to it.
But it's weirder to claim that Googlebot is aggressively DDoSing you.
And regardless of the cause, someone at Google has messed up by either letting themselves be used or screwing up the crawling algorithms to create so much load against one fairly minor blog.
That's hardly a DOS attack. I doubt if anyone would even notice, except that he has his server set up to alert on every 404. Sure sounds like paranoia to me.
That's potentially a lot of money wasted on nothing.
> My blog logs an alert when someone accesses a bad URL. The reason for the alert: a series of bad URL requests is often an indicator of an attack, or a precursor to an attack.
And he doesn't say it's using 25% of his available resources -- he says 25% of visits to his blog are from Googlebot. Since he doesn't mention what his total visits are, we have no way of knowing how much money is wasted, but I'd be willing to bet the amount is pretty close to $0.00.
- Is he sure it's Googlebot and not somebody spoofing headers?
- Is he sure Google is generating the queries, as opposed to Google just following URL's coming from elsewhere?
- Is he sure it's the search crawler, and not some browser service trying to cache autocomplete combinations or results?
- Has he tried simple potential fixes like changing GET to POST for searches, or robots.txt to set a crawl interval to something like 30s?
It just seems so weird. Google crawlers have finite resources so this seems like a misunderstanding or really bizarre edge case or both.
It's not clear if he was getting hit by GoogleBot or e.g. FeedFetcher, which does NOT obey robots.txt, because it's triggered by human actions: https://support.google.com/webmasters/answer/178852?hl=en
There's also the possibility, as you say, that traffic was spoofed. Google publishes its netblocks over DNS, so it's easy to double check that.
In the author's previous post on the subject [0], he verified that the request were coming from Google IPs.
> Is he sure Google is generating the queries, as opposed to Google just following URL's coming from elsewhere?
From the URLs quoted in the article, it looks more like some sort of (machine learning-y?) system trying to infer his site's canonical URL scheme.
[0] https://www.hackerfactor.com/blog/index.php?/archives/484-Go...
The main issues and complaints we run into are mostly unusual requests for indexing content.
People see the total number of requests and think it's a significant burden and just have an irrational need to protect their content.
It's almost a form of paranoia at some point.
I mean we have NO problem removing someone that doesn't want to be indexed and 99/100 we do so without a problem.
The only times we push back are when government organizations contact us as we feel we have a legal right to index the content. We haven't had to really fight over this because they usually agree and that's the end of the discussion.
Edit: add other; ambiguous.
User-agent: ia_archiver
Allow: /
User-agent: ScoutJet
Disallow: /blog
User-Agent: Googlebot
Allow: /
Disallow: /blog/index.php?/archives/2018/
Disallow: /blog/index.php?/archives/2017/
Disallow: /blog/index.php?/archives/2016/
Disallow: /blog/index.php?/archives/2015/
Disallow: /blog/index.php?/archives/2014/
Disallow: /blog/index.php?/archives/2013/
Disallow: /blog/index.php?/archives/2012/
Disallow: /blog/index.php?/archives/2011/
Disallow: /blog/index.php?/archives/2010/
Disallow: /blog/index.php?/archives/2009/
Disallow: /blog/index.php?/archives/2008/
Disallow: /blog/index.php?/archives/2007/
Disallow: /blog/index.php?/archives/2006/
Disallow: /blog/index.php?/archives/2005/
Disallow: /blog/index.php?/archives/2004/
Disallow: /blog/index.php?/archives/2003/
Disallow: /blog/index.php?/archives/2002/
Disallow: /blog/index.php?/archives/2001/
Disallow: /blog/index.php?/archives/2000/
#Disallow: /blog/index.php?/categories/
User-agent: *
Disallow: /blog
Disallow: /badbot-a2d0ac98abcaf3cafd9eff83b3cffa98fec7a390a6c5b9
I'm not 100% certain about that ? in there, but I think you can do this instead: Disallow: /blog/index.php?/archives/*/I wouldn't call it an attack, but it seems more like some sort of autocompletion checking given how similar it is to the queries generated to Google if you use their search autocomplete feature --- one request for each character entered.
That said, Google's search quality has already taken a steep decline and I've been getting blocked very often for searching more obscure things now, so whatever "bot mitigations" they put in, I hope they don't make that even worse.
Me too! It seems like nearly any time I go past the second page for results they think I'm a bot.
- As a user, I type "foo.com" into my address bar. Having already gone to a specific page of that site before, autocomplete "helpfully" expands it to that page's URL "https://foo.com/this/is/a/page"
- It turns out I don't actually want that, I want the main page, so I start deleting characters until I'm left with "https://foo.com"
If Google took each of those deletions as a separate URL (either by bug or on purpose) and tried to verify said URL is valid, I could see a GoogleBot request being generated for each iteration between the initial URL and the full URL, starting with that initial one.
So apparently not.
Is there any way to do that in robots.txt? Someone did mention adjusting the crawl rate in Webmaster Tools -- would the author need to do this for every search engine?
Sounds like his issue is mainly with Google, not other search engines, so it’s not an O(n) problem he’s got to solve.
/blog/index.php?/categories/14-Forensics (wanted)
/blog/index.php?/categories/14-Forensic (unwanted)
/blog/index.php?/categories/14-Forensi (unwanted)
/blog/index.php?/categories/14-Forens (unwanted)
How do you propose blocking the last 3, without having an enormous robots.txt?He removed the search before using a simple robots.txt...
I imagine what happened to this guy is that somebody coded up a site that used his as a service and google spidered it and went crazy trying to follow the links. It sucks, but that is what robots.txt is for. Anything that requires more than just serving a page should be blocked.
You can drive yourself insane reading your logs trying to guess what bots are trying to achieve. They seem to excel at spidering exactly what you don't want while ignoring the content that would actually be useful.
The one time I did have a problem with a misbehaving crawler was with something called 80legs[1]
[0] https://webmasters.googleblog.com/2011/11/get-post-and-safel...
[1] https://sheep.horse/2012/8/80legs_is_a_pain_in_the_neck.html
I'm sure their ML is pretty decent... but I do have to wonder if it's ever wound up leaving hundreds of comments on an anonymous forum or something, thinking it was a search box.
Google does field trials all the way in Chrome. Why would it be so unrealistic that they did similar trials with Googlebot?
Google bot is stupid af and their support is even less helpful.
TL;DR: Google just scans you site even if it always gets 404s for millions of requests.
That sounds insane. Is that actually true or is there an exaplanation of some sort?
As other comments mentioned, he could have easily prevented this with robots.txt:
User-agent: *
Disallow: /searchhttps://support.google.com/webmasters/answer/80553?hl=en
Then use robots.txt on any dynamic route like search.
A solid site map could help too.