We can do better than DuckDuckGo
drewdevault.com
drewdevault.com
Do they really though, for normal people that is?. Some of my searches today below, can't remember the exact terms I used. Mix of DDG and Google.
1) Walt Whitman, I wanted a basic overview of his work to satisfy some idle curiosity. DDG gave me his wikipedia page. Bingo
2) EAN-13 check digit. First result wikipedia telling me how to calculate it. I see it is simple and I have a long list in Excel to check. I can't be bothered to think so...
3) EAN-13 Excel. First result has an example that I copied and pasted.
4) Timezone [niche cloud system]. Said system didn't do what we expected, seems to be timezone issue. First article is discussing this niche issue and offers solutions
5) Does Shopify support x payments. Yes it does
6) Coronavirus test. Got straight to government site.
7) MacOS version numbers. First hit...
8) How come my Microsoft x platform is showing as being at y level of service when my Buddies is not. Straight in
Am I just a perfect search customer? I don't seem to be getting the problems Drew is?
I just searched for "lockdown rules for SA" (I'm in South Australia and we just had a new 20 person cluster, so we are going back into lockdown).
On DDG the first results was a Guardian article which was good, but then the rest were a mix of South African articles and blog spam. There were no SA Gov pages on the first page of results.
On Google the first result was the South Australian gov site with the rules, the second was the Guardian article, then more SA Gov pages and at result 8 I got a South African result.
https://searchengineland.com/google-reaffirms-15-searches-ne...
So at least they are trying to to location based results.
BTW, when I perform the same search, Google's first result is "What Are the Lockdown Rules for South Africa? A Guide for ..." and all the other results on the first page are about South Africa too. (Note: I'm in Japan)
Everyone in Australia uses "SA" - this is one of the reasons why location based context is important.
> You don't need the "for" either.
I worked on consumer search for a few years, and on text based search word like "for" are helpful to get exact match. Even if the term frequency of "for" on its own isn't particularly useful "for SA" absolutely is. (And these days with neural ranking using sub-word parts it is even more useful).
Searching for "lockdown rules" "sa" (together) just now gave a bunch of South Australian specific results with the "Australia" localisation setting enabled.
With the localisation setting disabled, all the results were indeed about South Africa instead.
It gets tiring quickly and I find easier to append !g instead of clicking the regional toggle button.
Sure, give me local results if I don't specify anything. But let me tell you if I want results in English now.
You don't have to. DuckDuckGo allows [0] to create a link with settings: https://duckduckgo.com/?kl=jp-jp.
www.google.com/ncr
'ncr' here stands for no country recognition. It allowed many expats to do technical searches without the noise of regionalization results.
Of course someone clever at google figured out that was probably too useful and now it just redirects you back to google.com because screw all those niche use-cases.
When you were in a different country (e.g., India), and you typed in google.com out of habit, it would recognize your IP-geo and redirect you to the country-specific domain (e.g., google.co.in).
If you really just wanted google.com for whatever reason, then you'd type google.com/ncr. It then wouldn't redirect you based on your IP-geo, and you'd stay on google.com.
In other words, google.com/ncr _always_ redirected you back to google.com. Then, and now.
However you can see from the comments in both android police [0] and reddit [1] that, irrespective of your assertiveness, the behaviour did indeed change at least in 2017 if not more times before.
It at the very least used to preserve the suffix and absolutely respect no regional results. It's the same as the old bolean operators, google claims the behaviour is unchanged but will silently ignore them.
[0] https://www.androidpolice.com/2017/10/27/changing-googles-do...
[1] https://www.reddit.com/r/google/comments/4xda1p/googlecomncr...
Google automates, DDG leaves me to choose. I prefer the 2nd approach every time.
This is exactly why I like DDG way more than Google and why I love to use Alfred instead of Spotlight on my Mac. With DDG you have !bangs and with Alfred you also can tell him what you’re looking for. 99.9% of the time I know I’m looking for a file or a folder or a definition of a word or want to open an app or want to search the web etc. With Spotlight you’re stuck to the order Apple designed the results to show up
I also search in Spanish (Castilian), but sometimes I want results from Latin America, sometimes from Spain.
Being able to set the language/region is of incredible help in both cases. There is no way to automatically detect this.
In DDG it's usually the first page.
I still do a lot of !g when I search technical stuff, as it lets +word -word and DDG doesnt find a lot of weird github issue pages, old forums, usenet posts sometimes.
Google does infer purpose better, and if someone is looking to buy something, it does well there too.
Ddg is very good at info queries and the more one uses it, the better it is.
What they could do is exactly what google did and that's to review those uses and improve.
But what they have right now is solid, given just a tiny bit of work.
If anyone is thinking of making the switch, you can always redirect your searches to Google by throwing a g! in the query.
I wonder if this is generational or cultural?
Personally, I dislike trying to interface with a machine using natural language, because I know it can’t really understand me, and I’d rather read and interpret the results for myself than have an algorithm pick the “best”.
I actually find speaking to machines (e.g. automated phone systems, Siri etc) using natural language quite embarrassing, as if we were pretending that real life was like Star Trek.
Until we address meaning, my statement remains solid. And it can often be easier to treat the tool like what it is rather than figure out how to best pretend it is something it is clearly not.
Although such queries are habit forming and now google does a decent job understanding the actual question
Got good at including words for context early on and never stopped.
My default remains word searches. Frankly, better operators would benefit me more than questions would.
I really do not always want to formulate a question. Doing that makes sense sometimes.
Often, I want to see relevant info, then formulate other queries.
I'm a bit lazy right now to remember all the problems it has, but some of the most obvious are looking up for news on recent events (especially something small, stuff that doesn't appear in reuters and these sorts of media) and trying to find out some basic stuff about local shops and such (of course, I only know about how it feels in my location, not worldwide). On both occasions I pretty much always use "!g ..." right away, because DDG is just clueless about this shit. Google does this just fine (in fact, sometimes it's even impressive: there are thousands of cities like mine, yet Google can often tell me where I can buy some stuff I'd have no idea where to look for).
This is exactly correct. Excluding poor local search results (which is understandable bc of the privacy aspect), Bing/DDG has trouble with long tale search query relevance (5+ word queries), and also finding results from small or obscure sites. The later is simply because Bing's organic index is not as large as Googles.
Bing/DDG's organic results are still very good, but they are not as good as Google's in the above specific circumstances.
My mother wanted to access Amazon last week and typed "my amazon account" in french in the windows search, which searched for those terms in Bing. One of the first (paid) results was a scam site triggering alarm sound, fake virus notifications and asking her to call a scam hotline.
At least DDG filters out the ads but the problem in this case is Bing's OS integration.
The last time I did a comparison, Bing did better (I don’t know what DDG does with the Bing results exactly, everyone says they just show Bing results, but no one knows and it simply doesn’t mesh with my experience).
Because Bing does not randomly filter out half my terms while DDG does even for "-forced terms. This is my #1 problem with DDG and I complain about it in pretty much every DDG thread (while otherwise loving DDG).
For few result searches, DDG shows you essentially random stuff even if they have the result I want (which can be tested by searching for an exact sentence from the result page). On the other hand doing the search on Bing gives me the result without neutering my query.
Other than that, I've been using [Runaroo](https://www.runnaroo.com) as well.
Now, even with quotes, I routinely get a whole first page of results where my terms are not included anywhere. Google generally respect the quotes.
I hope that a verbatim search function will be restored in the future; I think it's an essential basic tool for a search engine, and without it the user can be left with the impression that the engine either doesn't understand what it is being asked to do, or that it is wilfully disregarding instructions because it thinks — often wrongly — that it has a better idea of what the user is searching for than the user does.
I thought the whole point was to improve over time, not get worse. :(
(No, I do not consider typing "!g" before any search that contains quotes a solution to the problem.)
Edit: typos.
For your use case, simply append !g to the DDG search and it will do a Google search instead.
It also sucks at retrieving very new information.
And I say this as someone who set DDG as default.
I mean, you do seem to be DDG’s ideal user. You searched for mostly technical issues, and a hot political issue.
Still, for really niche topics if search needed there is no way around Google. On the other hand Google is not the only way to explore the web, let alone auto-complete a url...
Not a big fan of this conclusion. Who chooses the white list, and why should I trust them? Is it democratically chosen? Just because a site is popular very clear does not mean it's trustworthy. Does it get vetted? by whom? Also, who's definition of trustworthy are we trusting?
If I want my blog to show up on your search engine, do I have to get it linked by one of those sites, or can I register with you? Will I be tier 1, or
It would be bad if those in the positions profit by "authorizing" who is good though.
You can have really biased technically terrible filters that for example put a site on level 4 because it is to new, to small and any number of other dumb SEO nonsense arguments. (The topic was not in the url! There was poor choice of text color!)
I think wikipedia has a lot of research to offer on what to do but also what not to do. Try getting to tier 2 edits on a popular article? It would take days to sort out the edits and construct by hand a tier 2 article.
I expected that to evolve to get more specificity but things went completely the otherway and we can't even specify a term is on a page reliably with Google now.
Similarly, I was all in on xhtml and semantics (like microformats) where you'd be able to search for "address: high street AND item:beer with price:<2" to find a cheap drink.
I imagine for a FOSS solution we would have to make configurable every separable ranking algo and the option to toggle them in groups as well as build cli like queries around them (with a gui)
I'm starting to see a picture now. In stead of wondering how to build a search engine we should just build things that are compatible. A bit like The output of your database is the input of my filter.
Take site search, it is easy to write specs for with tons of optional features and can easily outperform any crawler. Meta site search can produce similar output. Distributed diy cralwers can provide similar data.
Arguably top websites should not be indexed at all. They should provide their own search api.
The end user puts in a query and gets a bunch of results. It goes into a table with a colum for each unique property. The properties show up in the side bar to refine results (sorted by howmany results have the property) Clicking on one/filling out the field/setting a min max displays the results and sends out a new more specific query looking for those specific properties. New properties are obtained that way.
You could have user compiled lists of sites to show in search results.
Let the users pick the lists they want to see, and communities can create and distribute lists within themselves.
Could easily be implemented for any current search engine.
To a large extent, this is what you already do when you view a page of search results. Filter them based on your understand of what sites / results hold value.
On a public scale, you could make an argument for tighter integration/better privacy with the lists. For example:
Browser -----Request-to-SE-----> Search Engine
^ |
| Unfiltered Results (In YAML/JSON)
| |
| V
|--Desired Results------ Local Filtering/Rendering
On a private scale, if you are only crawling sites on the allow list than you have the possibility of being able to better maintain a local database of sites to show up on the search.Edit: Possibly this could be easier to use to set up distributed search as well, as each node could index a given list, and then distribute that list similarly to DNS. Don't really know how well that would work though, just an idea.
I'm thinking about playing around with this in my spare time, but that part seems the hardest to do.
It’d take a lot of cycles while small but if you get a network growing you could even have sub-networks with their own whitelist additions (and every user has a blacklist.)
If I never see a pinterest link, or one of those sites that just republishes stackoverflow answers unedited, I'd be fine with it.
Of course, these systems always get abused and some political or news site will end up on it.
I think what I'd say in defense is that we've misunderstood what search engines are useful for. They're really bad at helping us discover new things. Your blog might be awesome, but it's not going to be easy for a search engine to tell that it's awesome. It's going to have to compete with other blogs that also want views, some of whom are going to be better than yours at SEO, and so on.
What a search engine might be able to tell is that it's useful. Because what search engines are at least potentially good at is answering questions. You do that by having a list of known good sites to answer specific types of questions, and looking at the sites they link to. It's when you try to do both (index everything on the web and provide accurate answers to specific questions) that you end up failing to do either. For example this is the #2 result for "python f strings" on DDG[1]. It's total garbage, and, quoting the blog, "we can do better". (This result is also on page 1 for the same query on Google.)
What I believe ddevault is suggesting is that we make a search engine that does the only thing search engines are really good at, answering questions. You throw away the idea of indexing everything on the web, and therefore the possibility of "discovery". What that means is that in 2020 you need some other mechanism for discovering new sites, bloggers, and so on. Fortunately we do have some alternatives in that space.
To be clear, I don't know if I 100% buy this argument, but I think it's the general idea behind what's being suggested in this blog post.
I mean, that's basically the core of original Google pagerank, right? A "good" site linking to another site is what makes that other site some amount of "good" too, links from better sites carry more 'juice'. "good" is of course not just binary, but a quantitative weight.
I don't know to what extent that's still at the core of their relevancy rankings. I don't know how all those annoying spammy recipe blogs or content farms get to the top of the results either. I don't think it's because Google's engineers believe they are "good" results.
Relevancy ranking on web search is clearly a hard problem, mainly because so many authors are trying to game it, it's a feedback loop.
If Google can only do as google does despite pouring a whole lot of money into it, I don't see a reason to bet that better will be what's basically an over-simplified description of how Google started out doing it (and then evolved it because it wasn't good enough).
I hope the project of improving on DuckDuckGo is successful though, and some of the other proposals in the article sound promising to me.
In the future I would like to see an open source search engine 'paid for' using some combination of homomorphic encryption, blockchain and Tor-style technologies to trade bandwidth and processing power with its userbase in exchange for search results and other services, but I don't have the expertise to assess how feasible that might be.
It definitely seems that search engines can’t find new websites for people. Now they are just aggregating Q&A.
There are other options that may be better, but in general, very few people are looking for them.
Here's your "python f strings" query on Runnaroo:
https://www.runnaroo.com/search?term=python+f+strings
I'm the creator of Runnaroo.
So you need to speed it up, but it does look good for searching documentation at least.
There could be a community of Software developers running one instance of the OS search engine that focuses on programmers needs: documentation, vcs hosters, dev blogs, tech news websites and on topic blogs. Great if you need to search for how to solve a software issue, terrible if you need to figure out how long to cook spaghetti.
The lists of crawled pages i guess would be visible/searchable too, at least on instances that you could trust.
If you don't like the list, and the maintainers for whatever reason don't want to change it to your liking, feel free to use another instance or set one up yourself.
I find it helps when looking for obscure info on random topics.
In other words, this is precisely how a market functions.
Still, it's good to remember that it was once uncertain whether people should access the web using something a table of content (portal/directory) or something like an index (search engine). It seems the search engines won.
That's part of the point. There could be different search engines that run the same code, but have different sets of tier 1 domains that cater to different audiences. And if you have the resources, you could set up your own engine with a set of tier-1 domains that you chose.
It uses Gigablast which has a much more fair search result set more akin to search engines of the past!
This is what we need more than anything. More independent blogs. The ability to search events now, or 10 years ago, mass indexing of RSS feeds, etc.
A general search engine is kinda way out of the ballpark for now. But you could specialize for long form blogs, from all sides, hard-left, hard-right, women in tech, white supremacists, all the extremes and moderates.
I've love to have an interface to search a topic and see what all kinds of people have posted long form, without commentary or Twitter/Facebook bullshit "Fact checking" notices. I what to see what real writers are seeing across the spectrum on a given topic for the week or month.
Thought experiment: what would a search engine look like if it only indexed RSS and Atom feeds?
The problem is that content farms have mastered the art of writing like an ostensibly independent blog. This is most visible in recipe blogs, where for example the site will look independent, the blog owner’s "About Me" page will say that she is a young woman born and raised in Louisiana and passionate about her home region’s cooking, but the English is replete with the sort of mistakes that non-native speakers make. You can tell that the content writing was farmed out to someone from Eastern Europe or Southeast Asia, and basically the whole blog and its owner are fake. (Even all the recipes were drawn from other blogs, but someone was paid to rewrite them slightly.)
You’re right that the walled gardens have hurt this. So often I search something specific, or a topic, and find very little. But I know there are communities on Facebook for this, I know there would be peoples posts out there on Instagram which 100% answer my question. But they may as well not exist. Unless I was “following” then when it was said, and mentally indexed it, these things are mostly unfindable, and that’s if I even have an account for said service (which I don’t for Facebook)
It’s sad, more people than ever using the internet, more content & knowledge being created than ever before, yet it’s no longer possible to find the great answers.
They wrote their own search engine.
They closed shop earlier this year.
If your goal is "to make something better than the Duck" and you succeed, the Duck dies... what is your goal now?
Ideally, services like theirs would be continuously audited by respectable, trusted organizations like the EFF.. multiple such organizations even.
Then I'd have at least some reason to believe their claims of not collecting data about me.
As it stands, I only have their word for it.. which in this day and age is pretty worthless.
That said, I'd still much rather use DDG, who at least pay lip service to privacy, than sites like Google or Facebook, who are openly contemptuous of it.
At the very least it sends a message to these organizations that privacy is still valued, and they'd lose out by not trying to accommodate the privacy needs of their users to some extent.
What I do care about is trust-building and monopolistic practices.
That, to me, is a great reason to use DDG instead of Google or even Bing.
DDG has been my default search engine for years and its results are good enough for me 95% of the time. I only need to use Google as a fallback when searching for niche technical information or "needles in haystacks".
Being super financially successful off free products and services is not a recipe for an honest, citizen respecting company.
[1] https://help.duckduckgo.com/duckduckgo-help-pages/results/du...
[2] https://help.duckduckgo.com/duckduckgo-help-pages/results/so...
This refers to the actual search results and if they used their bot for that I don't see why they would say "multiple partners" instead of "multiple partners and our own crawler". The fact that they don't have their own index is such a common "complaint", and this page is often referred, so if they really used their own bot they should have added that a long time ago.
And it's not just a legacy page they have forgot to update. It keeps being updated. In 2019 it said:
"We also of course have more traditional links in the search results, which we also source from a variety of partners, including Oath (formerly Yahoo) and Bing."
It's interesting to note that back in 2014 the page looked like this: https://web.archive.org/web/20131202065705/https://duck.co/h...
Here they talk about their own indexes getting bigger but at the same time admitting that "it seems silly to compete on crawling and, besides, we do not have the money to do so". Completely understandable but also interesting that the current page doesn't mention their own index at all. Maybe they used to have a goal to build their own independent index that has now been dropped?
All in all, I think it's safe to presume that their own crawler is only used for Instant Answers etc since that's the part of the sources where it's mentioned. Or at the very least used to such a small extent in the actual search results that it would be disingenuous to even mention it as a source.
A service, running on somebody else's machine, is essntially closed.
I think the only way to have an 'open' service is to have it managed like a co-op, where the users all have access to deployment logs or other such transparency.
Even then, it requires implicit trust in whomever has the authorization to access the servers.
But, I agree with you - and I don't think the author had really thought through what they were demanding, they made no mention of licensing other than singing happy praises of FOSS as if that would magically mean you could trust what a search engine was doing.
You mean AGPL https://en.m.wikipedia.org/wiki/Affero_General_Public_Licens...
Why would I trust someone to do that, though?
I think the next step forward should be to have indices that can be shared/sold for use with local mode. So you might buy specialised indices for particular fields, or general ones like what Google has. The size of Google's index is measured in petabytes, so a normal person would still not have the capability to run something like that locally.
Edit: In another thread, ddorian43 has pointed out the existence of Common Crawl,[2] which provides Web crawl data for free. I have no idea if it can be integrated with YaCy, but it is there.
You get a subscription and the index updates (enclosure+preloaded drive) are send to you periodically.
The front of the enclosure says: 2020 4th quarter
What am I doing wrong (or right), here? I put a thing in and find it. I just don't use Google any more.
Genuinely curious why it's working for me and such garbage for everyone else.
Honestly we shouldn't be using Google for everything. Why not just search StackExchange or Github issues directly for known bug problems? If you need a movie, !imdb or !rt forward you to exactly where you want to really search on.
If DDG or Google also included independent small blogs for movie results, I could see the value in that. I'd prefer someone's review on their own site or video channel, but it doesn't. We've kinda lost that part of the Internet.
Would be nice if they could do better on queries like that... though funny thing is if they didn't respect privacy they probably could. Log any searches where a user looked for something and then tried the same thing prefixed with !g. Use those for figuring out where to focus efforts and what to test with.
I always send feedback when I come across incorrect results and also try to when I get a really easy find.
I have not had to resort to any other search engine for at least five years.
I can think of some improvements (better forum/mailing list coverage), but it's generally pretty good. Lately if I don't find it on DDG I probably won't have much luck anywhere else, either.
For example, you might search for `vue js on show` whereas `vue on show` will show you (in the UK) results for what is on at Vue cinemas.
With Google, I expect it would understand that you are probably searching for JS related vue questions and rank those higher.
Then anyone who wants to use the data can either copy it to their own S3 buckets to pay just once, or can use it with some sort of pay-as-you-go method. Anyone who runs a search engine can use the algorithms as a guide for the specific searches they are interested in for their site, or can just make their own.
You could trust the other indexers not to give you bad data, because you'd have some sort of legal agreement and technical standards that would ensure that they couldn't/wouldn't "poison the well" somehow with the data they provide. Further, if a bad actor was providing faulty data, the other actors would notice and kick them out of the group or just stop using their data.
It would have to be fully open source, I agree with the other parts of Drew's essay here, but I think we could share the index/data somehow if we got together and tried to think about it. We just need a standard for how we share the data.
I don't personally think any system specification is impossible unless it goes against since mathematical law, so a really fast, distributed query system where there are a few hundred specialized providers for a single query is feasible. Imagine the aggregator does initial analysis to determine the category of search, like programming, news, or restaurant reviews, then sends the users query to a set of specialized providers that supply an index for that category, then fuse the results with some further analysis of the metadata returned. Then the user can also include or exclude the specialized providers at will.
You could also eliminate the aggregator as a service as simply make it a user application on the desktop, allowing for even more user control and maybe caching or something.
Though it would require development of the "charge to download" S3 buckets and infrastructure to support payments.
There is also an economic issue where you have to calculate the download cost to also cover storage costs.
Not sure I buy the example that is given here.
1. It's an issue in their browser app, not their search service.
2. It's not completely indefensible: it allows fetching favicons (potentially) much faster, since they're cached, and they promise that the favicon service is 100% anonymous anyway.
3. They responded to user feedback and switched to fetching favicons locally, so this is no longer an issue. https://github.com/duckduckgo/Android/issues/527#issuecommen...
> The search results suck! The authoritative sources for anything I want to find are almost always buried beneath 2-5 results from content scrapers and blogspam. This is also true of other search engines like Google.
This part is kinda funny because "DuckDuckGo sucks, it's just as bad as Google" is ... not the sort of complaint you normally hear about an alternative search engine, nor does it really connect with any of the normal reasons people consider alternative search engines.
That said, I agree with this point. Both DDG and Google seem to be losing the spam war, from what I can tell. And the diagnosis is a good one too: the problem with modern search engines is that they're not opinionated / biased enough!
> Crucially, I would not have it crawling the entire web from the outset. Instead, it should crawl a whitelist of domains, or “tier 1” domains. These would be the limited mainly to authoritative or high-quality sources for their respective specializations, and would be weighed upwards in search results. Pages that these sites link to would be crawled as well, and given tier 2 status, recursively up to an arbitrary N tiers.
This is, obviously, very different from the modern search engine paradigm where domains are treated neutrally at the outset, and then they "learn" weights from how often they get linked and so on. (I'm not sure whether it's possible to make these opinionated decisions in an open source way, but it seems like obviously the right way to go for higher quality results.) Some kind of logic like "For Python programming queries, docs.python.org and then StackExchange are the tier 1 sources" seems to be the kind of hard-coded information that would vastly improve my experience trying to look things up on DuckDuckGo.
I'd also love to be able to specify I want results from the last year without having to set it everytime.
As I understand it, you'd want to continue to search the whole "unbiased" web, then apply different filters / weights on every search. I really do like the idea, but I imagine we'd be talking about an increase in compute requirements of several orders of magnitude for each search as a result.
Maybe something like this could be made a paid feature, with a certain set of reasonable filters / weights made the default.
Maybe this would require too much data to be sent to the client, compared to the usual case where it only needs a page of results at a time. If so, would a compromise be viable, whereby the client receives the top X results and filters those?
I got signed up for goodreads (book review site), and I get tons of spam. It's not quite the same as your idea, but it is a currated list. I don't know how you stop spammers from adding bogus links in the python interest list (to use an example).
This is a hard problem..
EDIT: Clarified goodreads reference!
pro-privacy does not sit well with terms such as search history and user profile
But a system such as I'm describing is probably the only one that can be entirely consistent with the two disparate requirements of fully anonymizing users, and being useful to both programmers and ophiologists studying different things called "python".
It would basically be like sending a search query in this form:
"Python importerror help --prefs={mit_cs_club, studioghiblifans_new, britains_best_baking_prefs, AlpineMountaineersIntl}"
If you like baking, anime, and mountaineering, it's probably convenient to leave all those active for your searches, even your purely programming-focused searches. But you could toggle some of them off if articles about "helping to protect imported mountain pythons" are interfering with your search results, or if you want to be more anonymous. If you're especially paranoid you could even throw in a bunch of random preferences that don't affect your query but do throw off attempts to profile you. You could pretty easily write a script that salts every search with a few extra random preference lists, for privacy or just for fun, and make that an additional feature. The tool doesn't need to maintain any history of your past activity to cater to your search, so I think it would be a good thing for privacy overall.
the more anonymous among us tend to opt for common IP addresses and common user agents to become the tree among the forest. adding a profile to that would, well, only add to a digital fingerprinting profile
> those preferences don't need to be attached to you, just your query
that's not how it works. preferences are by their nature personal. every transaction would have your interests and hobbies embedded, on top of metadata
> mit_cs_club, studioghiblifans_new, britains_best_baking_prefs, AlpineMountaineersIntl
you have voluntarily made yourself the birch among the ebony
> If you're especially paranoid you could even throw in a bunch of random preferences that don't affect your query but do throw off attempts to profile you
how would they not affect the query? they complement the query. or rather, unnecessarily accompany the query. your results depend on your input. it doesn't matter what colour glove you wear to pull the trigger if you bury the gun with the body
> write a script that salts every search with a few extra random preference lists, for privacy or just for fun
just the latter. fuzzing would be pointless since the engine will have already identified you by now
it sounds like an annoying browser extension at best. to label it a pro-privacy tool would be ludicrous
The problem is that we have no say in ranking and filtering. I think it should be customisable both on a personal and community level. We need a way to filter out the crap and surface the good parts on all these sites. I am sure Google wouldn't like to lose control of ranking and filtering, but we can't trust a single company with such an essential function of our society, and we can't force a single editorial view on everyone.
As we have many newspapers, each with its own editorial views, we need multiple search engine curators as well.
As such, they effectively have a list of "tier 1" domains.
Any system that ranks things purely based on votes or view counts can have a feedback loop that can amplify "bad" results that happen to get near the top for whatever reason. For web search, this would encourage results that look right from the results page, even if they're not actually a good result of what the user is looking for.
An example of this would be when you're trying to find an answer to a specific question like "How do I do X when Y?". The best result I'd hope for is a page that answers the question (or a close enough question to be applicable), while the promising-looking-but-actually-bad result is a page where someone asks the exact same question but there are no answers.
I think this is a place where Google has pretty obvious algorithm problems. For example, I’m building a personal website for the first time in many years, and obviously that means I’m doing a fair bit of looking up new or forgotten webdev stuffs. It’s widely known that W3Schools is low quality/high clickbait/has a long history of gaming the SEO system. They’ve been penalized by Google’s algorithm rule changes but continue to get the top result (or even the top 3-5 results!), even with Google having a profile of my browsing habits, and knowing that I intentionally spend longer on these searches to pick a result from MDN or whatever. It seems pretty likely that W3Schools is just riding click rate to stay at the top. And it’s pathological.
for some languages, W3schools is as good a reference or better than the official documentation.
And they're definitely better than most seospam.
Edit: I semi-intentionally forgot to mention Stack Overflow because it’s so unpredictable in terms of quality.
> Second, we measure engagement of specific events on the page (e.g. when a misspelling message is displayed, and when it is clicked). This allows us to run experiments where we can test different misspelling messages and use CTR (click through rate) to determine the message's efficacy. If you are looking at network requests, these are the ones going to the one-pixel image at improving.duckduckgo.com. These requests are anonymous and the information is used only by us to improve our products.
The Firefox network logger does show requests to this domain when I click on a link in the search results, before the page navigates away. This suggests to me they might by logging this information. To be clear, this is speculation on my part, because I haven't examined the URL parameters in detail.
In any case, I'm not sure how much this manages to improve the results, since usually I can get help with my Python query (for example) using whatever crappy blog post is first in the results, but results from the official docs or StackExchange are still probably better and should be prioritized.
So I think that stepping back and re-thinking what a search engine fundamentally is, is a great starting point for disruption.
Additionally, something the OP didn't mention is that ML technologies have progressed dramatically since 1998, and that much of that progress has been done in the open. I can't imagine that not being a force-multiplier for any upstart in this domain.
There is a DuckDuckGoBot and I think it was an interview or podcast Gabriel did a while back that he mentioned they use it for filling out gaps in the Bing API data to provide the instant answers, favicons. Their preference for the instant answers were authoritative references such as docs.python.org. This would have been a while back though.
With that logic, Apple’s OCSP server is also 100% anonymous (which I legitimately can believe it is).
The problem with this strategy is always going to be that different users will regard different sources as most desirable.
For example, it's enormously frustrating that searching for almost anything Python-related on DDG seems to return lots of random blog posts but hardly ever shows the official Python docs near the top. I don't personally think the official Python docs are ideally presented, but they're almost certainly more useful to me at that time than some random blog that happens to mention an API call I'm looking up.
On the other hand, I would gladly have an option in a search engine to hide the entire Stack Exchange network by default. The signal/noise ratio has been so bad for a long time that I would prefer to remove them from my search experience entirely rather than prioritise them. YMMV, of course. (Which is my point.)
I knew close to nothing about building a company or a project, or how a proper business model would have helped it. I was the leader (SABDFL) of the group, and unfortunately I didn't lead it well enough to succeed. We had some good ideas, but ultimately we failed at building more than the initial prototype.
The idea behind it was simple: WorkerBee nodes (users' computers) would crawl the web, and provide the computational power to run Beeseek. Users could upvote pages (using "trackers" that anonymously "spy" the user in order to find new pages - repeat: anonymously). The entire DB would be hosted across multiple nodes. Auth and other functionalities would be provided by "higher level" nodes (QueenBee nodes). Everything was going to be open source.
Well, it didn't work.
Thankfully, because of Beeseek, I met a few very smart people that I am in touch with to this day.
Life is strange and beautiful in its own way.
Weird, though, that today I still believe that Beeseek could have been the right thing to build. Who knows?
One difference from what you describe is that the OP is specifically recommending against decentralization/federation, where it seems to have been the core differentiator of your effort. I don't think what OP is describing is quite what you are describing.
Most people still believe that it's possible for one search engine to help anyone find anything without it knowing anything about them, which is just ridiculous. To get good search results you practically have to read someone's mind. Google basically does this (along with their e-mails, and voicemails, and texts, and web searches, and AMP links, and PageRanked crawls, and context-aware filters) and they still don't always get it right.
There is no magic algorithm that replaces statistical analysis of a large corpus along with a massive database of customized rulesets.
> We should also prepare the software to boldly lead the way on new internet standards. Crawling and indexing non-HTTP data sources (Gemini? Man pages? Linux distribution repositories?), supporting non-traditional network stacks (Tor? Yggdrasil? cjdns?) and third-party name systems (OpenNIC?), and anything else we could leverage our influence to give a leg up on.
Oh, great, so become the Devil himself, then. Count me out.
Microsoft tried and failed to build a competitor and it's not like they have shallow pockets.
They grossly underestimated a number of aspects:
- The huge number of man-years invested in Google's search quality stack hand-tuning and what it would take to replicate it.
- The fact that the machine learning field was simply not ready to tackle the search quality problem.
- The infrastructure required to build a crawler / indexer stack as good as Google's
I think in 2020, the second problem is within reach of many companies technically. It's mostly a matter of throwing enough money at optimized infrastructure.However, replicating the search quality stack is going to be very hard, unless someone makes a huge breakthrough in machine learning / language modeling / language understanding at a thousandth of the cost it currently takes to run something like GPT-3.
The most likely candidate to execute properly on that last bit is - unfortunately - Google.
I don't see it that hopeless. I feel it kinda is like starting Open Street Maps. It won't be perfect for a long time but there will be people who'd prefer it and help out.
Of course they would - it would be set as the default search on their iPhones with no clear-cut way to change it. You know, "security". The users don't know what's best for them, etc. as Apple seems to think.
We’re already talking about building a search engine, might as well talk about a model to convincingly detect blogspam too.
The recent UK Competitions and Market Authority report evaluating Google and the UK search market came to the conclusion a new entrant would require about 18 Billion GBP in capital to become a credible alternative search engine, in terms of size, quality, hardware, man hours making it.
Remember Cuil? Had the size, the fanfare but unfortunately not the quality.
Google pays Apple more than that every year just to set Google as the default search engine on iPhones.
In a way, Google is funding its future competitor.
Here's my brief and slightly made up history of search engines:
In the beginning of time, search engines took a Boolean query (duck AND pond) and found all the documents which contained both words using an inverted index and then returned them in something like descending date order. But for queries which had big result sets, this order wasn't very useful and so search engines began letting users enter more "natural language" queries (duck pond) and sorting documents based on the number of terms that overlap with the query. They came up with a bunch of relevance formulas - tfidf, BM25 - that tried to model the query overlap. But it turns out this is tricky because user intent is a really tricky problem and so modern day search engines just declare that relevance is whatever users click on. Specifically they just model the probability that you're going to click on a link (or something) using a DNN that uses things like the individual term overlap, the number of users that have clicked on this link, the probability it's spam, the PageRank etc. Some search engines like Google also include personalized features like the number of times you have clicked on this particular domain - because for instance as a programmer your query of (Java) might have different intent than your grandmother's. This score then gets used to sort the results into a ranked list. This is why search engines (DDG included) collect all this data - because it makes the relevance problem tractable at web scale.
Maybe just my perspective but I just really don't understand why OP would want to build an index - it's hard boring expensive and doesn't violate data privacy - and I don't think people grasp that - at least to some extent - data privacy and relevance are in direct conflict?
dmoz looks pretty great but the categories look limiting.
https://github.com/jmqd/folklore.dev
It's not even really at the POC stage yet, but I hope to host it with a simple web frontend sometime soon. Primarily, this is just for myself... I just want a good way to search the sources that I myself trust.
- "100% of the software would be free software, and third parties would be encouraged to set up their own installations" - I'm planning on open sourcing it under AGPL soon, once I've got documentation, testing etc. ready. Plus it's easy to set up your own installation (git clone; mkdirs for data; docker-compose up -d).
- "I would not have it crawling the entire web from the outset" - That's one of the key features of my approach, only crawling submitted domains. I'm focussing on personal websites and independent websites at the moment, primarily because I don't currently have the money for infra to crawl big but useful sites like wikipedia, but there's nothing to stop people setting up their own instances for other types of site.
- "who’s going to pay for it? Advertisements or paid results are not going to fly" - A tough anti-advert stance is another key differentiating feature to try to keep out spam, e.g. I detect adverts on indexed pages and make sure those pages are heavily downranked. Planning to pay running costs via a listing fee, which gives access to additional features like greater control over indexing (e.g. being able to trigger reindexing on demand).
For example, I decide Google is terrible when I'm searching for product reviews, and all I get are results to Amazon referral websites and spam blogs that never owned the products to begin with. So, I find 200 sites or forums that actually have quality reviews and I create a whitelist of those URLs, and I name it "John Doe's product reviews list".
Other people visit the search engine and they can see my list, favorite it, and apply it to their results. I maintain the list, so they continue to get updates as it's refined.
The idea is you visit the search engine, type your query, then select from a drop down one of your favorite curated lists to apply. Maybe you like to use "Mike's favorite free stock photo websites" when searching for free photos for your projects. Maybe you like to apply "Jane's vegan friendly results" when searching recipes or face creams. Maybe you want to buy local, so you use the "Handmade in X" list when searching for your next belt. Maybe you use a list that only shows results from forums, or another for tracking/ad free websites.
Keep track of list changes. So, if someone gets paid off to allow certain sites on their popular list, others can easily fork a past version of the list.
If you didn't start at the very early stage of tiny web (e.g., Google in 1996 as a research project) and grew with the web over the past 20+ years, or you don't have super deep pocket (e.g., Microsoft Bing in mid 2000s), then it's almost impossible to build a decent web search engine within a few years.
It's possible to build vertical search engines on far smaller scale, far less complex, far less lucrative things that Google/Microsoft has little interest today (e.g., recipes [2], podcasts [3], gifs [4]...)
It's also possible to come up with a different discovery mechanism for web (or a small portion of web), other than a traditional complete web search engine. Essentially you don't cross moat to attack a huge castle (e.g., Google). Instead, you bypass the castle [1], as it becomes irrelevant.
[1] https://twitter.com/benedictevans/status/1038538688232226817...
I know there have been talks of set-ups that essentially take a web archive of your entire history to search back through...
PeARS was meant to be installed voluntarily by users who would then choose to share their indexes only to those they personally trusted, so the idea is very privacy conscious but also very hard to scale.
Cliqz, on the other hand, apparently tried to work around that issue by having their add-on bundled by default in some Firefox installations[5] which was obviously very controversial because of its privacy and user consent implications.
I still think the idea has potential, though, even if it's in a more limited scope.
[1] https://github.com/PeARSearch/PeARS-orchard [2] https://cliqz.com/en/whycliqz/human-web [3] https://blog.mozilla.org/press-uk/2016/06/22/mozilla-gives-3... [4] https://blog.mozilla.org/press-uk/2016/08/23/mozilla-makes-s... [5] https://www.zdnet.com/article/firefox-tests-cliqz-engine-whi...
I was aware Mozilla had some involvement with Cliqz, but didn't really pay attention, I remembered the company became owner of the Firefox Addon Ghostery some years ago. They closed shop mid 2020, but their tech-blog 0x65.dev is still up. There are a lot of posts from last December that explain its inner workings.
User-agents really do contribute their history, containing which search terms led to which pages. From this data (they named it the "human web") the search engine had a page model of which search terms led in higher frequency to the page. Related search terms were normalized. Only later did a "fetcher" really index high frequency content and consider it again in a later stage of search. Interesting bootstrap, more energy efficient maybe as it can run on less information.
Sending the search and browsing history offsite needs explaining and trust. But ultimately, any centralized search engine will see the search data too. Cliqz approach was trying to piggy-back on the result sets on search terms by established search engines, the search term + choosen result combination a donation of the user. Not any less invasive then other search engines. Would I send off the whole corpus of my browsing history? this raises good questions. Thanks for the links!
I like this attitude. Makes me happy to be a paying member of SourceHut.
If AdWords targeting was purely based on the search term, I don’t mind.
The search engine has to generate revenue somehow, and the revenue generated on “Saas crm” with a single click is likely to be larger than any users annual subscription. (10 - 100+ per click)
I’m unclear on the ethical / privacy concerns of “AdWords” style advertising.
Agree, keyword (and location, a lot of searches are for X near me) for the most part offers a way of delivering relevant ads.
Google are able to generate more income per search because of their critical mass of searches and advertisers, as well as having more data on searchers based on search history to maximise that revenue per search.
A search engine's job is to present you with the best possible results for any given query.
A ad is either A) the best possible result or B) not the best possible result. If the ad is the best possible result, then the search engine must display it anyway in order to fulfill its mission. If it is not the best possible result, the search engine must violate its mission in order to display it. To put it bluntly, advertising is paying to decrease the quality of search results.
One way to make search better would be to embrace bias and give people what they want. Just accept that most information is biased and bring it to the forefront. Initially there would be some default domain whitelisting, but users can request sites to add or remove from their bubble. Maybe they can "share" these lists among each other and certain clusters would form, which you can then use to recommend more nodes. Users would also be clearly told that results are biased based on their preferences, and that they can choose to view results from other points of reference. It would also have different "modes" for things like products, information, and news. Maybe some users only want to see results from independent or ad-free sources, and they can choose those clusters. Maybe they want only far-right or far-left sources, and then they can choose those. At least people would be aware of their bias, which I think will help fight it more than acting like it doesn't exist. Essentially there would be an explorable graph so I can see different realities. It would be a combination of search and social media.
Maybe there could be some pure anonymous ad-free search engine but it more realistic to have alternative commercial one. I really dont care that people are looking at my searches for how to resize an array or cheap hotels in Florida.
Not to mention the cost, not sure something like this could be sustained with a Wikipedia-esque "please donate $15" fundraising model.
[1] https://searchengineland.com/google-reaffirms-15-searches-ne...
[2] https://www.weforum.org/agenda/2019/09/chart-of-the-day-how-...
The open alternative to Google doesn't need to have the same capacity for download and indexing.
I always start with DDG and revert back to Google if it doesn't help, or I feel "there's got to be a better way".
That said, talk is cheap, show us your engine.
I didn't know this. I tried to use DDG for a while after I switched to Brave, but the results were just not very good and completely missing results I was looking for at times. This would explain it if it was coming from Bing.
Trust me, there is not a ton of potential "just sitting on the floor" in web search.
There's no reason you couldn't allow the first N number of api hits to be free, then charge for higher tiers of access.
When you search for "kittens" you get the links that are most upvoted by the community.
If nobody has ever submitted links for search term "kittens" , you get a link to selected generic search engines. And "kittens" end up into a list of words someone has searched but nobody has yet added a good result link for.
I mean, you could always put a big red badge on top of the results that says something in the line of "this search term seems to be troublesome. You may want to check qwant/ddg or maybe even google."
A) DDG is already better than Google for some search queries
B) I use Epic Search and that is good, though it just uses Yahoo or Bing results (not sure which)
C) We should have a lot of search engines
It should not be Google paying $8 billion a year to Apple and $X a year to Firefox/Mozilla and leveraging Google Search as Default in Chrome
and monopolizing search
You have one of two options. The crowd-funded approach would have to come with an understanding that you're trying to build a better, more private search engine that will leave the payer paying for all those who don't and the payer won't be able to have a say in anything as we'd like search results to be flat and even across the board. That means if I search for Coronavirus results, not only do I get the goverment and "Official" sources, I should get everything that I'm looking for and refine as needed.
The second approach is obviously big money but if you have big money coming in, big money will give You one of two options; Do as they say or they withdraw funding leaving you back at option one and having to downsize.
Rocks and hard places people. Not much else you can do there. Unless you take a Pilled.net approach.
The beauty of such oss solution maybe the custom heuristics that can be created based off the crawled data.
You can block crawlers if you can identify them, but reliably identifying them is hard.
After all, Google uses 'Mozilla' in their Google Bot user agent String for similar reasons - because sites might expect it.
Running an instance for websites related to your occupation or hobby YaCy is quite wonderful. You don't want google removing a bunch of pages that might cover exactly the sub-topic you are looking for. Of course the smaller the number of pages in your niche the better it works.
I like this idea. It would be interesting to see the domains of every search query that I have clicked on and see what the distributions is like. I suspect there would be a long tail but I wonder how many domains actually need to be indexed for 99% of my personal search needs. Does anyone have data like this?
My suggestion? Older folks don't have a well known search engine targeted at them, and there are features that could make a search engine helpful for them (High contrast mode or built in screen reading for the vision impaired, anti fraud results for common scams that target seniors based on search results, links to places to watch old shows when people search for character names), and they are a lucrative demographic, both in terms of having money to buy things but also for political ads.
I'm sure a focus group of seniors would have more detailed thoughts.
How would an open-source search engine stand against abusive SEO-optimization? If anyone can understand how the ranking algorithm works then anyone can game it.
You can also create such bookmark manually and use %s in the url as a placeholder where search query should be placed.
The manual configuration can be useful when there's no direct search field. For example freshports.org allows querying freebsd.org. I can add a bookmark with search keyword "fp" to point to https://freshports.org/%S
After that I can type in address bar: fp lang/python39 to land on https://freshports.org/lang/python39 (the capital %S doesn't escape special characters like /)
(1) actually based on https://github.com/jivesearch/jivesearch/tree/master/bangs
Now _that_ is putting your money where your mouth is!
Glad to see a technology leader taking this important issue head-on.
The harder question is "who and why is going to pay for development?"
Open source means you can't sell ads, any sane person would search for a fork without ads.
That approach will not scale in the general case, not even with the search engine following links from said sites. Too many areas of interest, too many languages, nowhere near enough people to categorize 'high-quality' sites all the time.
It could work for domain-specific searches for expert communities. SourceHut should start with a good code search engine...
I don't agree with that at all. But if your goal is to make "a better search engine" as you said, it does actually have to be "better" and not just different.
The one that strikes me most is the results. I feel like DDG doesn't search the entire internet, like there's zillions of pages there waiting to be indexed, even old websites, but the results i get are so poor.
Even with this handicap, I still use over INSERT YOUR AD HERE Google Search.
If the first page doesn't satisfy, just prefix the search terms with '!b ' and Bing usually nails it.
For example: https://news.ycombinator.com/item?id=25093009
it would basically make already famous domains shine and dump lesser known domains into 20th page oblivion.
it saddens me because google search results actually helped you discover new sites and new people, but it's been years since that has changed.
> ⇒ This article is also available on gemini.
Does anyone know what is gemini? I couldn't find relevant results in any of the search engines.
Google is like an oracle because it exploit all of this. And it works, with a cost for your privacy (but maybe you are ok to pay it). DDG is more a "meta-search-engine" with limited capabilities. Bit you got the flexibility of accessing the search engines of thousand of websites.
If you don't care for your privacy, don't use DDG and stay with Google.
They have their own crawler too, beside aggregating results from Bingo and others.
Every spammer will have a test harness for the SE stack to ensure their spam is ranked as highly as possible.
If it can be made to work, it should continue to work.
No on point search results from anyone, as far as I can tell. I suppose DuckDuckGo is in a suburb, called Paoli.
devault weinberg philadelphia
Nothing useful for that either.
(might be because google's results these days are so bad though... can't really tell)
We can't have a search engine that is only useful for finding the most relevant web pages for a given query. People love highly relevant advertisement in their search results.
Proceeds to give many reasons why it would be very difficult to do better.
1. Account deletion
2. GDPR data request
3. Option to unsubscribe from emails
So right now his blog reminds me one famous US politician Twitter account. Never fix your own problems, just blame others more often.
From wikipedia:-
"Fascism (/ˈfæʃɪzəm/) is a form of far-right, authoritarian ultranationalism characterized by dictatorial power, forcible suppression of opposition and strong regimentation of society and of the economy which came to prominence in early 20th-century Europe. The first fascist movements emerged in Italy during World War I, before spreading to other European countries. Opposed to liberalism, Marxism, and anarchism, fascism is placed on the far right within the traditional left–right spectrum."
If it walks like a duck, and quacks like a duck... it is, at the least, a precursor to fascism. It doesn't happen all at once, and if we wait until it's plainly obvious to call it like it is, then it'll probably be too late. And even if you don't want to call it facist, it's out of line to compare this shit to my personal blog. That's a baseless character attack, which is a dick move, not to mention against the HN guidelines.
How to get quality results, and a sustainable, community-led search engine?
=== Contexts ===
A "search engine" such as Google is good at many things, and extremely bad at others. The main issue with it in my view is that it lacks context about what you are looking for. The main context you can ask for is "Videos", "Pictures", etc.
* Specifying context takes time, so it's OK for long searchs Google sucks at (find this specific article I read a while back). Take your time while you specify language, exact/fuzzy match, publication date, background color, author name or any number of things you know about your search.
* Some requests can be processed with instant answers, that's good news as it fits the open source model quite well.
* Lastly, the other requests. Some are asked like a question and might require NLP to sort trough. Quite hard IMO, it might get better but will still require compute power if done server-side. It's mostly: parse the question to find the context, and perform a contextual keyword search/instant answer.
* And those that aren't questions: "regular", keyword-based requests, that "just" require a big index and a big infrastructure to search it.
=== Hardware ===
Now, we are left with the cost centers: hardware. IMO, the only way to scale is to rely on the community and distribute things.
* Databases: if this is a community project, and not too latency-sensitive, the community can help by distributing them over a p2p network, even with a single source of trust.
* Queries, walking the database: delegating processing to untrusted third-parties is a bit more dangerous. Maybe allow each user to specify a list of trusted servers? Can be centralized and clients ask the network, though it might leak part of their search, depending on the index method. Could be client-side?
* Processing the answers: client-side, or trough any number of frontends (like searx).
* Crawlers: crawling the net isn't cheap. You could use one or multiple sources of trust. Domain-specific crawlers, like hinted at in Drew's post. Maybe crawl on demand or trough the user's computer (web extension that indexes as the user browses, and allows them to full-text search their history; share it or not).
=== Content ===
For some measure of quality, I find that websites that do not have advertisements offer better-quality content. That's likely due to conflicting interests. It would be great if the semantic web mandated disclosing revenue sources. You could downrank or avoid crawling sites with ads and/or Google Analytics, for instance. This could be abused to an extent if the service ever becomes popular, but heh https://xkcd.com/810/
Domain-specific crawlers would be nice as well.
=== Added value ===
To be adopted, the service needs to be better than the original in some ways. I think that a new search engine should not try to conquer the masses at first. Instead, find some people that are not satisfied with the current offering and court them. Currently, I think this isn't met by advanced search: exclude websites protected by recaptcha, only include websites that are less than X years old, no ads, etc.
Allow users to create their own contexts and easily switch them: bangs, tabs, date/time/geoip, etc. Have them create contexts dedicated to their activities: programming is an obvious one, but so is cooking, gardening, encyclopedic search, language usage/dictionaries, etc.
=== Monetization ===
At that point, I am not sure it can ever turn a profit? EU grants? Consulting? Help webmaster set up search on their own website? Sell desktop indexing software?
Well, I do have some ideas around content curation, but I am a tad reticent to share them here, and not sure they are more useful than the above.
I found out through comments on hn that 8chan was backup under a new name: 8kun
Typing it into google I get articles about it but no link in the results.
In duckduckgo first link.
Made me think what else am I missing?
Recently I found myself desperate for any information on a price of hardware i had gotten. I was swapping out all sorts of queries woth different keywords hoping to find a manual. I was able to find some marketing material which was helpful, albeit barely. Eventually I had exhausted the search results for most pf my queries, gave up and assumed that it was simply lost to time and I was out of luck.
Eventually I went back to the sales paper I found. Going to the site it was hosted on, a Lithuanian reseller. I translated the page, eventually finding a direct link to a user manual on the exact same page as the sales paper I had found. The document was in English, contained important words from my queries (such as the product name, company, "user manual" etc. The document was at the same path as the sales paper too. I hace no idea why Google found the sales paper but not the manual.
Unfortunately the manual still wasn't what I was looking for exactly but it was a hell of a lot better than what I could get from Google's results.
"We also save searches, but again, not in a personally identifiable way, as we do not store IP addresses or unique User agent strings. We use aggregate, non-personal search data to improve things like misspellings."
So they save your web searches and claim that they do so in an non-personally identifiable way. The privacy problems with this claim are many, even if one accepts it at face value (good luck verifying that this is the case).
Every centralized search engine has immensely hard-to-resist and powerful incentives to play "The Eye of Sauron" with your data. Additionally, they offer single points of compromise to other, far more powerful actors. Whatever guarantees DuckDuckGo gives you -and right now they don't give any- don't mean much, if they've been thoroughly (willingly or unwillingly) compromised.
Which doesn't mean one should always steer well clear just that one should at least be aware of the tradeoffs one makes when using a centralized search engine. And with DuckDuckGo's misleading marketing, I feel that this point is lost on significant chunks of its userbase.
It wouldn't matter anyway, because decentralization doesn't really solve privacy any better than centralized search, besides the fact that it could theoretically provide more choices.
No matter what you use, privacy ultimately depends on trust. The reason that I have more trust for DDG than I do Google is, unlike Google, its primary audience is privacy-minded folks. If it came out that DDG was tracking users and selling that data, DDG would be immediately done as a brand. They at least have some incentive to do what they say. Decentralization provides no such benefit because a search "node" is unlikely to have any sort of meaningful brand to keep up.
> And with DuckDuckGo's misleading marketing, I feel that this point is lost on significant chunks of its userbase.
How is it misleading? My understanding from their marketing is that they don't create profiles of their users based on searches. Until we have evidence to the contrary, it's not outrageous to assume they are being truthful.
In your Privacy page (Data Usage Section) there is a mention of stored "Browser Data" & " These logs contain the time of visit, page requested, possibly referral data, and located in a separate log browser information." & "We may also use aggregate, non-personal search data to improve our results".
This is an honest question - How is that not exactly what the Parent stated was the issue?
So they save your web searches and claim that they do so in an non-personally identifiable way.But agreed that all search engines have to be trusted on their word about anonymising data and not retaining PII when it comes to searches specifically. There's nothing any front end user can do to verify it.
Okay, can you list just a few?
If you're going to make counter-claims like this, you're going to have to provide evidence.
Statements like these are not conducive in gaining popular support for increased privacy.
How do you verify that DuckDuckGo does -the minimal and ineffective- things they claim to do? They offer no proof.
How do you verify that DuckDuckGo does not secretly cooperate with more powerful coercive actors?
How do you verify that DuckDuckGo, offering a single point of compromise, has not been thoroughly compromised by more powerful actors?
Save a sha256 hash of every search for 24 hours. If you see the same hash from >10 distinct IP addresses in a 24 hour period, save the search terms.
That's just off the top of my head, I have no reason to think they're doing it exactly like that. The point is that you're claiming that we shouldn't trust DuckDuckGo because you can't think of a way that they could securely and privately do what they do -- but that's just your intuitions, for whatever they may be worth.
I also don't really buy the worries you have with the last two questions, e.g.:
> How do you verify that DuckDuckGo does not secretly cooperate with more powerful coercive actors?
How would you verify that for any centralized service, open source or not? I think your security concerns go a bit beyond what most people interested in critiquing / improving DDG can reasonably expect to achieve.
Other centralized (search) services don't have their entire existence depending on this one factor. What is DDG if not alleged privacy? Just use Bing directly.
I think it's entirely reasonable to be in the following posture: I want as much privacy for my web searches as I can reasonably achieve without having to run a search engine myself. I'm willing to trust that search providers are not saving personally identifiable information or passively turning over search data to law enforcement if they claim that they are not in their terms of service.
That's pretty much the use case for DDG. With Bing you know they are violating your privacy. With DDG you have a promise in writing that they are not. It's hard to see how that's not strictly better than what you get from Bing if privacy is among your core desiderata.
>I'm willing to trust that search providers are not saving personally identifiable information or passively turning over search data to law enforcement if they claim that they are not in their terms of service.
Do other search companies disclose that they share data with the FBI, NSA, etc in their ToS? Genuinely don't know.
I think, technically, some sort of honeypot verification could prove a compromise (i.e. if information that has very little chance of existing naturally in two systems, say a string a guids).
But... I agree with your point. I don't think this is actually feasible or realistic, just technically possible.
To a first approximation, you just... do it.
Granted, if you search "{jerf's realname here} {embarrassing disease} cure" or something, in the pathological case, you could at least guess that maybe it was me, though even then my real name is far from unique, and nothing stops anyone else from running such a search.
But otherwise, if all you have is a pile of a few billion searches, you don't have any information about any of the specific searchers. Even if you search for your own specific address, you don't really get anything out of it; there's no guarantee it was you, or a friend of yours, or an automated address scraper. There isn't much you can get out of a search string without more information connected to it.
The rest of your criticisms are too powerful for the topic at hand; they don't prove we shouldn't use DDG, they prove we shouldn't use the internet at all.
A search history can reveal many things about a person. The mere fact that someone, somewhere searched for "star wars harry potter crossover slash", unconnected to any other search item, doesn't reveal anything about anybody.