MixRank (YC S11) Raises $1.5M from Mark Cuban, 500 Startups
techcrunch.com
techcrunch.com
We're working on interesting problems at scale, like crawling the web and big data analytics. If this sounds interesting, it's worth reaching out :).
That's strange. Crawling Google's ads is prohibited by robots.txt and their terms of service. And Google tends to enforce this kind of rules.
On one hand you'd think if this wasn't permissible, they'd be big targets and Google would go after them legally. On the other hand, what would the legal basis for blocking this be?
Someone browsing my website and seeing a Google ad has not agreed to any Google terms of service. It's a tough argument that a browser is allowed to do this but a crawler isn't, especially when Google's own crawlers now include javascript-executing webkit browsers.
They could enforce it by technical means, by just blocking your IP.
If this were the case, then you wouldn't be able to run a headless browser without accidentally violating TOS and robots.txt.