The Wrong Way to Get Noticed by YC
mattmazur.com
mattmazur.com
Might be a good way for HN to avoid getting hammered by crawlers and still let the hacker-types slurp the data.
It's actually kind of hard to upset him. Frankly, he has a lot of money. It's going to take more than one errant script to ruin his day.
http://www.google.com/search?q=cache:news.ycombinator.com/it...
At this moment, that's the latest cached item they have. So, it'll lag by a few days, but that wouldn't matter much here.
"We're sorry...
... but your query looks similar to automated requests from a computer virus or spyware application. To protect our users, we can't process your request right now.
We'll restore your access as quickly as possible, so try again soon. In the meantime, if you suspect that your computer or network has been infected, you might want to run a virus checker or spyware remover to make sure that your systems are free of viruses and other spurious software.
If you're continually receiving this error, you may be able to resolve the problem by deleting your Google cookie and revisiting Google. For browser-specific instructions, please consult your browser's online support center.
If your entire network is affected, more information is available in the Google Web Search Help Center.
We apologize for the inconvenience, and hope we'll see you again on Google. To continue searching, please type the characters you see below:"
when i check it in my browser it was fine, but as soon as i started scanning with the software google blocked it. they may have recognized that the requests werent coming from a standard browser, so they flagged my ip.
I crawl Google frequently and that format has always worked for me... At least it did until I made this post. :)
Google has a lot of conflicts with media companies and it always starts out with google ignoring some "rules" expecting to either win in court or settle at some point.
So I think breaking these rules is part of the process of how sensible rules are established in the first place. Yes it's recursive, but HN readers should be smart enough to understand that ;-)
Why not? Google obeys robots.txt -- if you don't want them to crawl your site, it is trivial to arrange. I think violating Google's terms of use is pretty hard to justify as ethical.
And it will only take a few hours to update later.
PG had the email address from that first message. How he made the link between the server load and Matt's index is the real question.
Well done. :)
How is that interesting?