curl -I http://digg.com/robots.txt
HTTP/1.1 404 Not Found
Content-length: 1443
Content-Type: text/html
Date: Wed, 20 Mar 2013 16:27:13 GMT
Server: nginx/1.1.19
Connection: keep-alive
an HTTP 404 robots.txt should not be an issue - but maybe there was a robots.txt with something else in there instead, and now it's gone.but even so, if there was an all blocking robots.txt site:digg.com would still show the URLs of digg.com, as crawling is optional for indexing. (maybe they used the non-documented Noindex: / directive, but i doubt it)
if you want to achieve such a clean removal, in most cases you must request a complete removal via google webmaster tools.
so yeah, somewhere, someone might have screwed up, most likely on diggs side, maybe on googles side, maybe a combination.
UPDATe: now it's
curl -i http://digg.com/robots.txt
HTTP/1.1 200 OK
Cache-Control: public
Content-Type: text/plain
Date: Wed, 20 Mar 2013 17:05:29 GMT
Etag: "c47ccf1a49c24cc5842430aa75c72ef491292412"
Last-Modified: Wed, 20 Mar 2013 16:51:48 GMT
Server: TornadoServer/2.2
Content-Length: 24
Connection: keep-alive
User-agent: *
Disallow:
which is the best practice robots.txt (allow all, or rather: hey, i have a valid robots.txt file but the second statement is not a valid robots.txt disallow directive, so this means allow all as there is no other disallow directive (note: i once coded https://npmjs.org/package/robotstxt which tried to reimplement googles spec of the robots.txt https://developers.google.com/webmasters/control-crawl-index... so i have some experience in reading robots.txt files ))still, it wasn't a robots.txt issue
hmmm, just a hypothesis: maybe just maybe someone thought it was a good idea to remove www.digg.com via google webmaster tools from google (as their main domain is digg.com and not www.digg.com and they definitely tried had some www.digg.com URL indexed, as even bing some some of these URLs indexed http://www.bing.com/search?q=site%3Awww.digg.com&go=&... ) but google is (maybe) set to treat www.digg.com and digg.com as the same (via the settings panel in google webmaster tools), removing www.digg.com then could result into removing digg.com as well (something similar happened to a client of mine years ago), so could be the issue, could not be the issue, we would need more data (access to GWT) to verify this.
User-agent: *
Disallow:
This robots.txt should allow all bots to search the entire website. However, I think Google also penalizes websites that serve different content to googlebot than to non-bot user agents. Sitemap: http://blog.digg.com/sitemap-pages.xml
Sitemap: http://blog.digg.com/sitemap1.xml
User-agent: *
Disallow: /private
Disallow: /random
Disallow: /day
Crawl-delay: 1 User-agent: *
Disallow: User-agent: *
Disallow: