The day I was DoSed by Google
thekeywordgeek.blogspot.com
thekeywordgeek.blogspot.com
- Use rel=nofollow on links you don't need to have followed (this prevents passing of PageRank, which generally means we're less likely to crawl them)
- Use 503 for rate-limiting crawlers. 503 means we'll just retry later.
- Use the crawl rate limit in Webmaster Tools (I see you submitted the report there, so that should be active soon)
- If the content is fully auto-generated, you might choose to use a "noindex,nofollow" robots meta tag on these pages to prevent them from being indexed separately. It's hard for me to judge how useful your content would be in search directly.
A 503 would still require a GAE instance to be running so wouldn't necessarily deal with my problem.
I have seen "noindex nofollow" kill a site stone dead in the past so I am very wary indeed of using it. In my experience once you've noindexed a page it is nigh-on impossible to get the engine to index it again.
My content is autogenerated, though I hope it has enough value to be considered useful. It's time-series data of word frequencies in politics, so for example you might use it to see how one candidate is doing relative to another in an election campaign.
http://www.languagespy.com/politics/uk/trends/70th/70th-anni... http://www.languagespy.com/politics/uk/trends/70th/70th-anni... http://www.languagespy.com/politics/uk/trends/70th-anniversa...
I can't check at the moment, but my guess is that all of these generate the same content (and that you could add even more versions of those keywords in the path too). These were found through crawling, so somewhere within your site you're linking to them, and they're returning valid content, so we keep crawling deeper. That's essentially a normal bug worth fixing regardless of how you handle the rest.
And persistence to track how many crawl requests have been served in the last N minutes. Even blindly serving a million 503's an hour could get really expensive.
[edit] That being said, the issue would be the same if it was another hoster or another search engine. I guess the real solution would be to be able to limit the crawl rate, as the OP said.
So, while I agree with the sentiment that it sucks that this crawling eats the quota, the solution is not to simply bypass the quota.
I'm a bit confused. What computation does the GoogleBot cause to be performed that benefits the Google service user? (Not Googlebot related stuff like indexing).
EDIT: Thanks kyrra!
How does not charging for outgoing network traffic make computation free? You'd still be paying for everything else, eg the instances themselves, datastore storage, read/write datastore calls, using the logs API, which means mining bitcoins wouldn't be free.
If requests initiated by Googlebot were free to run, you could make a giant website full of garbage and use each free request to spend 50ms mining bitcoin.
Otherwise you'd have a team at every cloud provider trying to figure out how to manage bots.
That, and adding proper support for tuning the crawl rate.
The more you think about this, the more insanely complex it gets.
I no-indexed all of them because it had thin content. Guess what happened to my traffic? Almost no effect.
Stop assuming Google will send you traffic for auto generated pages - do you really think Google will even display them in the first few pages over quality content that actually is written by human beings?
Allowing those auto generated pages to be indexed will do you more harm than good. Noindex them.
I'd probably go for not allowing spiders to crawl more than a few chosen pages (home, about, etc) until you have enough revenue to support it going to other pages.
From https://plus.google.com/+PierreFar/posts/Gas8vjZ5fmB (Not sure how official this is but Pierre appears to work for Google)
Primarily the section
"2. Googlebot's crawling rate will drop when it sees a spike in 503 headers. This is unavoidable but as long as the blackout is only a transient event, it shouldn't cause any long-term problems and the crawl rate will recover fairly quickly to the pre-blackout rate. How fast depends on the site and it should be on the order of a few days."
Edit: Looks like the over quota page is a 503. Couldn't hurt to do it early yourself, Googlebot will see it the same way whatever provides it the 503
Or maybe even just block crawling of the entire site except for the homepage?
What I'm hearing is that you built a massive application, you've run into a technical problem and now you would rather wait on Google to fix it than to take any suggestions on how to get it up for actual users to use. Seriously, don't do this - at your stage, it would be better to have 10 real users than a site that has been fully indexed by Google.
On your note about persuading Google to index your site after being excluded, do you have any actual experience with this happening?? I've been doing this kind of stuff for years and years and have never had a problem. It can take five or six weeks at the outside, but that is still less of a problem than a product that can't be accessed...
You can also tell Google to index your site more slowly, in Google WebMaster Tools, although if I remember right the setting expires every few months, and needs to be reset.
The odd thing here IS that webmaster tools won't let him restrict the crawl rate, that's very odd.
(Also, it would be nice if you could restrict crawl rate in robots.txt, not just webmaster tools).
In the end though, if your business is going to depend on Google indexing it, then you don't really want to tell Google not to -- or, really even to tell it to index more slowly. But a robots.txt can be a temporary measure while you figure out what to do -- if you want Google to index your site, you've got to make your site able to stand up to googlebot traffic. Caching is often pretty helpful, and can help with your site's reliability and performance beyond googlebot issues.
That's kind of just the way it is, right? If you want google index, you've got to be able to handle googlebot. Nothing too shocking here?
Caching is definitely something to look into, that can improve the reliability and performance of your site beyond just dealing with googlebot.
I guess the odd thing is that Webmaster Tools is not letting the author rate-limit. And it would be really nice if google defined and respected some extension to robots.txt to do it there. I guess you could always rate-limit google bot with your own firewall-ish tools, but it might make googlebot mad and you might get even less indexing than you wanted.
Really, if your product's success depends on google indexing it, you don't want to slow it down anyway, except maybe as a temporary measure -- you're going to have to figure out how to handle it. People are usually complaining about how to make sure googlebot comes to _more_ of their site _more often_, not the reverse!
(edit) Yes, the bot is still hitting the GAE site atm even though it's returning a quota error.
In a nutshell, if you put up millions of pages and tell google about it it will index you, if you don't want that you'll have to make choices about the quantity and/or switch to a different kind of host.
Also, this kind of 'bot trap' tends to attract penalties so if this is not some ploy to get traffic out of google you may want to re-consider how you've laid things out, the difference between a legitimate site with a lot of generated pages and a page-spammer is hard to determine and google tends to err on the side of caution.
Another option would be to move to a dedicated server. You can get quite a powerful server from a company like LiquidWeb for a few hundred dollars a month. (A "managed" server, so although a bit of know-how is needed to get it performing optimally, they can help you with the basics at least.) I expect with a bit of tuning of your web server (nginx or even apache with mpm_event or worker) you could handle that level of traffic even without caching, but you could also use something like Varnish to do even better.
Its source is the English language, so if there's a word or phrase that gets used, it has a result. Corpus linguistics is fun like that.
The cost effectiveness really depends on whether your data would fit into that 1TB or it'd require much more.
Saying "I have infinite data" is not an excuse for not looking for alternatives.
That's the messed up part. I guess the question is, does robots.txt override that or not? If it does, fine. If not, all you need to do is make a few "google ignores robots.txt" posts and the problem solves itself.
It's not really ignoring robots.txt either, as crawl delay is not an 'official' setting
Google would crawl our GAE site at bursts of about 30,000 requests in 4-minute periods. We had some quota exceeded moments.
On the other hand we got to load test our MongoDB backend in GCE without writing gatling tests. The results weren't promising for our ~$180/month VM.
..and so googlebot indexes ALL THE THINGS eating my quota, if only it indexed my garbage slower.'
Cool story brah.
Actually it is, google should detect its own GoogleBot and not charge people for using up all the traffic. Because this can be used on purpose to have people pay more. Very interesting artcile. Thank You, Jenny.
Especially if they haven't been there before.