The simplest thing to do is NOT parallelize your scraper, and sleep for a few seconds between requests.
Or do you also plan on suing Google once they scrape your site?
> "scraping" site every 15 minutes with massive ddos
Google doesn't scrape every 15 minutes. It was implied that they scraped the site often enough to cause a DDoS.
Additionally, it clearly wasn't a DDoS. Quoting from the parent:
> Some swedish equal rights server was "scraping" site every 15 minutes with massive ddos. We put capacha for the server ip only (let them scratch their heads now) and let it be
If there is only one server ip, it is by definition not distributed. Just a misuse of the term DDoS.
Google has a variable timing crawler. Google News crawls in almost realtime, but all Google crawlers back off if the site responds slow.
Additionally, Google respects `/robots.txt` (which allows the site owner to define a crawl delay) and uses sitemaps for hints, so it doesn't have to re-crawl every single HTTP object "every 15 minutes".
Did you mean "We couldn't handle the traffic caused by their scraping which lead to a denial of service"?