The DuckDuckBot crawls and indexes the web. http://duckduckgo.com/duckduckbot.html
The DuckDuckBot crawls and indexes the web. http://duckduckgo.com/duckduckbot.html
Every test search I've ever done on DDG shows near identical results to Bing.
I'd like to see a search that uses it's own data, examples?
Gigablast was the last serious third-party backend that had a chance for independent data. It's like old-school Google.
Gabriel should try to buy Gigablast and merge it with DDG so he has his own independent dataset.
Checking to see if it's hit any of our sites.
3k pages, not bad. Data is kinda stale though.
Interestingly, the fetches do not have a user-agent that identifies itself as the DDG crawler:
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322)
I'm assuming this is the crawler because it does not fetch anything besides text/html.
Gabriel, does DuckDuckGo's crawler have a distinct user agent? Can you talk more about how DuckDuckGo observes/respects robots.txt?
http://techzinglive.com/page/1028/179-tz-interview-gabriel-w...
around the 70 minute mark Gabriel mentions that DuckDuckBot is mostly about determining if pages are spam.
But that is a non-issue as far as I'm concerned. As long as the results are relevant and they got them legally, then who cares where or how they came from?