A method to use Google for DDoS. Bug or Not?
chr13.com
chr13.com
A simple fix will be just crawling the links without the request parameters so that we don’t have to suffer.
Many links would fail/have different content if the request parameters were removed from the URL. Perhaps the crawler could use some kind of reverse bloom filter [1] to be more careful/back off if it receives the same content from multiple URLs. However nothing is simple at Google scale so there are probably issues with this approach too.[1]: http://www.somethingsimilar.com/2012/05/21/the-opposite-of-a...
=image("http://targetname/1.jpg")
=image("http://targetname/2.jpg")
=image("http://targetname/3.jpg")The advantage of the querystring-method is that you can just find one suitable (i.e. huge) file and force Google to pull it down many times.
Maybe Google should consider putting a bandwidth limiter of some sort on that (or even better: use hashes to avoid duplicates), but I think screaming "security! vulnerability!" is not a good action to take here...
If we ignore that ETags are related to URLs and not 'files', ETag as suggested by userbinator might work for some cases, but if the large file is dynamically generated, it's unlikely to have an ETag; defaults in many servers are to make an ETag based on the inode of the file rather than any properties of the file, so if there are multiple servers behind a load balancer, they're likely to return different ETags.
(I know that servers can be configured not to send ETags or break caches by sending random ones every time, but this could reduce the data usage considerably since most of the responses would only include the headers.)
Rate limit per website (e.g. don't download more than 10 images per domain per second)
Limit the total number of images it downloads per document, so a single user can not cause too much traffic.
=image(CONCATINATE("http://example.com/?", RAND()))
If you add this to a spreadsheet and fill a few thousand rows with it. Each time the spreadsheet is loaded, google will hit the server a few thousand times."It's taking a while to calculate formulas. More results may appear shortly."
I set the spreadsheet document to load images like so: =image("http://example.com/image?id=146&r={increment here}")
After 30 or so images, google starts to slow down its fetch rate.
It wouldn't be too hard to block by User-Agent: Mozilla/5.0 (compatible) Feedfetcher-Google; (+http://www.google.com/feedfetcher.html); if you notice the traffic.
Feedfetcher does not fetch robots.txt though; so you'd have to do something in your server config.
[edit: fixed a typo, and agree with the update]
I would hope that Google is able to detect abuse of their infrastructure for (D)DOS.
I don't think removing the parameters would be ideal, though, since some sites might legitimately serve up different images based on different parameters.
Just limiting the amount of traffic to a single server, or outbound from a single spreadsheet, seems like a good solution, though.
You see a surprisingly high amount of excel-based bill templates - and you may want to hotlink the company logo or a signature.
Doesn't that somewhat defeat the purpose of a dynamic image?
Moreover, the issue is not about the feature, it's whether Google should limit the number of requests made per image. From other comments it seems like Google is hitting each image hundreds or even thousands of times. I suspect this is for cache? If that's the case, Google should look at better way to handle it. A single fetch and propagate to closest zone should be enough. But this is not a reason to limit the feature (eliminating parameter query).