An Interview with DuckDuckGo's founder, Gabriel Weinberg
techspot.com
techspot.com
I still think the real silver bullet will be to make a backend system that can peer with other systems to gather data and can be customized to get deep results from a small subset of the web that interests a particular user ( think corporation uses it for internal search and to make their site show up better on a common front end). I even started a research project to begin working with some of the things needed to accomplish it[1].
Trawl. The web gets trolled enough as it is.
troll ... 4.(intransitive, fishing, by extension) To fish using a line and bait or lures trailed behind a boat similarly to trawling; to lure fish with bait. [from circa 1600]
The first being speed. People complain about DDG's speed already. Relying on a host of external searches would only cause more issues. Especially if as Google reports that a huge amount of queries are unique.
Another would be ranking. Who determines what is the most relevant result to a query? While you have multiple sources knowing which one to go to has to be determined somewhere. If you hand it over to the peers, they can game the system by insisting they are the most relevant. If you leave it to the server what incentive do the peers have to participate? If you use an open algorithm which people use whats to stop people from gaming the system?
I don't think there is an easy answer to the search game. You either need to build on others systems to have something compelling, or have millions (if not billions) of cash to have enough runway to build/improve your own system, improve it to the point its worth using and then turn a profit.
I don't even think its the complexity of the problem that stops the second. It's bandwidth and disk storage. My personal prediction is once we get disk's up-to a size where you can store a sizable chunk of the web, coupled with enough bandwidth to crawl it in a reasonable time you will see more innovation in the search space as the barrier to entry will be lowered.
Google reports 20-25% of queries are unique, way more than I think most of us would expect. http://www.readwriteweb.com/archives/udi_manber_search_is_a_...
We can probably assume it's not evenly distributed across people. Some people would probably always get cached queries, while others would need fresh results at least half the time.
With that in mind, and with the tech world generally driving adoption of players in the search space I can't see the approach working. Nobody would switch when the queries are massively slower, even if they were 99% accurate.
If you look at http://donttrack.us/, it says your data "can potentially end up" in those places, not "will be sold to..."
The main thing that would really improve my experience would be if it was faster. Google really spoiled my with the instant results and suggestions as you type.
I'm also looking forward to them switching to SPDY as I always use the encrypted version, and this should make it a bit more responsive.
Another thing that annoys me is that they have ads near the top of the results that don't load quite at the same time as the results, so I sometimes am about to click on the top result but an ad pops up and pushes everything down making me miss my click. If they could somehow avoid that happening, it would also make the experience better.
- Some times the instant search simply doesn't work until I press enter (and sometimes, even when I do, the results just don't appear) - I noticed that google doesn't take you directly to the page. It takes you first to an intermediate page (probably for tracking purposes) and then redirects you to your result. I would be fine with that, but sometimes my browser just got stuck in the intermediate page (no idea why).
I sometimes need to use !g too and I agree that google sometimes is faster than DDG (at least when it works). What I really miss from google is the autocomplete
https://duckduckgo.com/?q=frequency+of+letters+in+The+quick+brown+fox+jumps+over+the+lazy+dog
https://duckduckgo.com/?q=days+between+6%2F22%2F1979+and+10%2F5%2F1979
https://duckduckgo.com/?q=hn+duckduckgo+interview
https://duckduckgo.com/?q=php+xml_parser_create+example
https://duckduckgo.com/?q=currently+in+theaters
https://duckduckgo.com/?q=msft
https://duckduckgo.com/?q=currency+in+panama
The only one I think Google does better is "currency in panama" however it also gets the information wrong in the "zero click" answer. The only reason I like that result more is the Wikipedia answer on the right is just more appealing to my eye.To digress a bit, zero-click is great when the information you want is actually accessible with zero clicks, but it's very, very limited: as soon as you need to click through to a website, special widgets can't compete with a solid backend for regular search results. That's why I can hardly imagine switching to DuckDuckGo...
Interesting both Google and DDG have the linked article as result number one for the term "duckduckgo interview" but no Zero click for either.
That's an interesting comment about the zero click. It's interesting, but I find myself using zero click info more these days. I tend to craft queries which I know will pull this information back for me. Its much the same way that Google trained me to use their syntax or how the instant search trained me. I agree that a solid index of the web is critical for regular search results though.
I would be curious to know if the apparent issues with Bing/DDG are more down to people perceiving Google to know the answer. I know that when I use alternative search engines I get weird looks from people at work who say quote "Why are you using that? Its crap! Just use Google."
A HNSearch (HNSearch.com) for "interview duckduckgo" doesn't return this thread at all, in any of the results (even on a 'stories' only search), however the one thing it does surface is your comment, because that exact phrase was found in it. I realized we weren't showing the right comments, and I found a very small bug which has now been fixed. So thanks for getting me to notice that :)
However, if you search "interview duckduckgo's" (ie. the same wording in the thread's title) the ONLY result returned is this exact thread (an HNSearch limitation).
These same results are fed to us by the HNSearch API and so the fault lies within HNSearch's search methodology (as far as I can tell). We're still looking for a resolution to this, however, any suggestions are welcome!
With the launch of DuckDuckHack it will be interesting to see what people build for the platform. Plus, I'm just excited to see a talented team take on search.
1) It's slightly slower than Google, which became more apparent after the fifth or sixth search of the day.
2) No images integrated with search results. I didn't realize how often I searched for images until I used DDG. At Google I usually get a few images and an "Images" link to click. At DDG I needed to add a !gi to my query.
Like JohnsonB said, there really need to be another draw besides just privacy.
In some way it's 'trained' me to perform more targeted searches from the outset.
The slight lag before getting results is noticeable though.
I'm going to stick with the upstart for the moment nevertheless.
Google meets my needs better and their recent power search class was very cool.
A little off topic but I have had more than a few fantasies about starting my own micro search engine. Text analytics and knowledge management in general have been an interest of mine since the early 1980s. What stops me is that if I were to invest my personal resources in this I would want tens of thousands of users getting value from my system every day, and frankly, I don't think I could achieve that. Gabriel gets an order of magnitude more than what I would hope for, so I hope that he is very satisfied with what he has achieved. Good job!
* http://help.duckduckgo.com/customer/portal/articles/216392-a... * https://github.com/duckduckgo/duckduckgo/wiki/DuckDuckGoPerl * http://www.perlmonks.org/?node_id=848999
They've also been visible in sponsoring the Perl community.