Google bots follow SQL-injected URLs from third party sites
blog.sucuri.net
blog.sucuri.net
Just secure your website as if the attacker is not going to use the Google bot as a proxy, because well, nobody can guarantee that the attacker will use the Google bot.
Google bot is nothing more than a computer sending HTTP GET request to your server, so if your site is vulnerable to HTTP GET requests why you write like there is something special about Google bot?
I just can't justify why do you expect Google to spend resources on not running bogus HTTP GET request when anybody can run those? What is different about being hacked by Google bot and being hacked by an unsuspected user who clicks on a bogus link that was put on the same page where Google found that link to your server? Just doesn't makes sense.
Not only that, but it seems to me that it'd be a more efficient use of resources to spend the time hardening your own site rather than lobbying Google to implement something that only mitigates one potential attack vector. Even then, it just seems stupid because I'm sure there are valid GET query strings that might have select, insert, update, delete, or some permutation thereof in them.
It seems to me that it's just a punt on poor programming habits...
http://en.wikipedia.org/wiki/Medireview#Blocked_emails
Even if you look for somewhat complete SQL strings, if I want to host http://try-sql-in-your-browser.io/?sql=select+foo+from+bar, I'd want Google to index it.
http://try-sql-in-your-browser.io/?evil=true
It works well for IPv4.
On a related note, if you would like to be the person looking through the crawler code and designing a defense for this, and are willing to work in the bay area. Send me an email :-)
Not looking for work atm, so my advice: this is not your problem.
http://not-a-real-host.com/some_page/
And that page has a link for comments which takes you to http://not-a-real-host.com/some_page/?show_comments=1
And the designer didn't neuter the show comments link and if you follow it, it takes you to: http://not-a-real-host.com/some_page/?show_comments=1&show_comments=1
Ad infiniteum. That is a bug but the same sort loop where the cgi arg is going from ?page=1, ?page=2 is valid.Basically folks who love working on web crawlers are people who like figuring out puzzles like this. The number of interesting things web sites can do that make crawling them not work is unlimited I believe. We have lots of regular checks (things like how many unique documents have I pulled from a host, or how many pages have the same MD5 hash, etc.) Designing defenses is part of the fun.
So any time the crawler can be a bit more aware about what it is heading into, the saves it time not looking at pages which we know apriori we would not be indexing anyway. And that (which pages should be indexed) can only be answered from the set of pages that are crawled.
> We are contacting Google about it, but it is always something important to keep in the back of your mind. You can’t just whitelist their IP, and allow through without any type of inspection.
I wouldn't be surprised if the authors knew of the sqli vector, and whitelists appropriate clients. The problem is the sqli, not WHO can hit it (or who was ultimately responsible).
Edit:
I just went and looked at the company behind this article. I was too quick to judge apparently. Their business is protecting developers from stupid mistakes like sqli at the firewall.
Therefore, the article really IS about indirection. Sorry.
They're doing it for various Typo3 versions (I know, because we got some false positives in the past - Google Webmaster tools warned us about it, we saw the requests in logs, our fault for replying with Status 200 for some invalid URLs where we just showed our main page), they might be doing it for other software where the only way to check for vulnerabilities is to try an actual (harmless) SQL inject.
That's a bit weak, and might not pick out decent targets. However, visit the forum as a user, mention some specific sites, and hey presto - google infects them for you, while you disappear in the crowd.
Seems a bit airport thriller though. More likely the leaked sqli and google's use of it were accidental.
This is a pretty interesting attack vector, though. The old Google bomb being used for more nefarious purposes than spoofing page rankings.
Whereas HTTP goes over TCP which is a connection-oriented protocol. TCP offers message integrity by going back and forth between client and server multiple times to verify the message was retrieved successfully. Without a valid source address the 3-Way TCP Handshake used to establish the connection cannot succeed.
I think you can still register and download the videos: