How to make fun of Google Bot (PicoLisp Wiki)
picolisp.com
picolisp.com
What did work, however, was fool a searchbot: The whole thing got me a very angry mail from a Dutch search engine team (ilse.nl) whose bot had been stuck on it for an entire day. I had no robots.txt (didn't even know what it was), which the search engine team decided was a really nasty case of lack of netiquette.
So, somehow their ill-coded bot crashing was your fault? It's not like you forced them to crawl your site.
http://code.google.com/web/controlcrawlindex/docs/robots_txt...
the robots.txt of http://picolisp.com (found at http://picolisp.com/robots.txt) allows the indexing of http://picolisp.com/21000 and all follow up pages.
why? see the spec:
The disallow directive specifies paths that must not be accessed by the designated crawlers. When no path is specified, the directive is ignored.
see here https://github.com/franzenzenhofer/robotstxt for coffeescript implementation of a robots.txt parser
User-Agent: *
Disallow: /21000/
Which is what you deserve for using non-standard URL formats.
When I say "non-standard", I am saying am saying that if the website's URLs looked like "/21000/foo" and "/21000/foo?page=2", it would have been easier to craft a "Disallow" rule that would have successfully blocked all of the desired pages.
User-Agent: *
Disallow: /21000
or User-Agent: *
Disallow: /Google has a free robots.txt checker that lets you test your robots.txt files. Given a robots.txt file, you can enter specific urls and check whether that url would be blocked or not. Here's a link for more info on that free tool: http://www.google.com/support/webmasters/bin/answer.py?hl=en...
robots.txt on ticker.picolisp.com says "Disallow: /", but ticker.picolisp.com redirects to picolisp.com/21000, and the robots.txt on picolisp.com says "Disallow:". If he wants Googlebot to stop crawling those URLs, he needs to add "Disallow: /21000" to picolisp.com.