It does not do any of it's own indexing.
It's just a frontend to other very very expensive backends that have millions of dollars behind them.
The entire company can be shut down overnight if it's data feeds are cut.
It does not do any of it's own indexing.
It's just a frontend to other very very expensive backends that have millions of dollars behind them.
The entire company can be shut down overnight if it's data feeds are cut.
The DuckDuckBot crawls and indexes the web. http://duckduckgo.com/duckduckbot.html
Every test search I've ever done on DDG shows near identical results to Bing.
I'd like to see a search that uses it's own data, examples?
Gigablast was the last serious third-party backend that had a chance for independent data. It's like old-school Google.
Gabriel should try to buy Gigablast and merge it with DDG so he has his own independent dataset.
Checking to see if it's hit any of our sites.
3k pages, not bad. Data is kinda stale though.
Interestingly, the fetches do not have a user-agent that identifies itself as the DDG crawler:
Mozilla/4.0 (compatible; MSIE 6.0; Windows NT 5.1; SV1; .NET CLR 1.1.4322)
I'm assuming this is the crawler because it does not fetch anything besides text/html.
Gabriel, does DuckDuckGo's crawler have a distinct user agent? Can you talk more about how DuckDuckGo observes/respects robots.txt?
http://techzinglive.com/page/1028/179-tz-interview-gabriel-w...
around the 70 minute mark Gabriel mentions that DuckDuckBot is mostly about determining if pages are spam.
But that is a non-issue as far as I'm concerned. As long as the results are relevant and they got them legally, then who cares where or how they came from?
Let me know if you find a search result that is different.
How often do you do single word searches in the realworld?
As for "followed by bing results", there is a single one that is equal to bing. The results that follow are also far from similar (this is exacerbated by Bing insisting on giving me local results regardless of relevance, which DDG purposedly avoids).
And yes, I do single word searches very often, but if you're so inclined, here are results for a longer search: http://cl.ly/image/0m1E3I3J0M1U (DDG shows Hulu, tv.com, CTV, amazon, and doesn't repeat the Wikipedia entry)
If you adjust the term to "hacker movie", you get more similar results between DDG and Bing. But, overall, it does seem like DDG is returning more differing results now than in the past.
Same argument can be made about Google, they do not produce their own data just copy webpages from content providers. If content providers decide to block them from indexing their websites, Google would be irrelevant.