Why robots.txt and favicon.ico are bad ideas and should be eliminated.
bitworking.org
bitworking.org
If there was a similarly simple and effective solution, either in 1994 or now, he would have suggested it. But he didn't. We have to guess that he'd want something involving a declared robots-rules-'link' elsewhere.
But layering robots-rules as a 'link' from the headers of the root page, or markup tags (if the root page happens to be HTML), still requires an initial investigative hit to a site -- and if at the '/' page, probably involves more bandwidth than a compact robots.txt or 404.
Assuming a hostname is a unified 'site' is not perfect -- but in 1994 and now, that's the only inherent and well-defined unit of administrative control provided by the HTTP protocol.
By going with this easy-to-understand, easy-to-implement convention, webmasters have had enough freedom to opt out, and crawlers enough freedom to collect, to enable ~15 years of dazzling growth in powerful search applications. And robots.txt files built to the 1994 standard still work, handling 99.999% of what webmasters need to communicate to crawlers.
That success satisfies my design aesthetics.
Not necessarily. If the location were specified in the http headers then a robot could use the HEAD command first. That's relatively minor bandwidth cost and still lets the robot get the site rules without first getting any content.
(It's possible to strain and devise a convention that uses less bandwidth; that's why I said 'probably'. But that would require even more complexity -- such as the server being smart enough to send each robot only the rules that apply to it.)
Configuring a server to emit special headers is also harder for most webmasters than dropping a text file into a conventional location. And a HEAD-then-robots process is more complicated for robots-writers.
For the purpose of warning off robots, the /robots.txt placement convention was a very, very good solution on many axes, including minimizing traffic and adoption costs.
Only a peculiar early-optimization based on a certain aesthetic sense -- and being concerned about being a bad example for other similar applications that aren't a pressing issue even now, 15 years later -- can justify Gregorio's opinion that /robots.txt "was a not-so-good idea when the robot exclusion protocol was rolled out".
And you're right, the current scheme is easy. But it doesn't cover all the cases and makes certain types of sites impossible to do correctly (see http://news.ycombinator.com/item?id=639396 for an example). Putting a link to the correct robots.txt for a URL in the HTTP headers would be way more flexible and only slightly raise the bandwidth cost.
You could make it backwards compatible by assuming the current /robots.txt path if there is no header--So it would cost nothing to current site operators. Spider writers would have to do one extra HEAD but I refuse to believe that's terribly difficult.
Moreover, his proposed replacement is incredibly browser-centric. It requires any compliant crawler to contain an HTML parser. And woe befall anyone who typos the META tag, or manages to confuse the HTML parser before it gets there! The robots.txt specification, by contrast, requires no such heavy lifting: just request a single fixed URL.
META tags make perfect sense for favicon and the like - I won't dispute that. Robots exclusion, however, is a special case - it belongs outside HTML, not inside it.
So it is semantically correct, although most modern SEs do not do this, to index a site / URL that is disallowed via robots.txt, using link data alone.
Nowadays domain names are cheap and I don't see as many sites like that any more...
http://bitworking.org/news/431/wave-first-thoughtsImplementing a meta tag on the index page could work, but why change the already-working system that we have?
<link rel="icon" type="mime/type" href="....">
But if you are using the default favicon.ico it needs to be an icon, and you can include multiple sizes (which you can't do with png, etc).