Building a Search Engine from Scratch
0x65.dev
0x65.dev
I'm skeptical that they'll be successful, but I wish them the best. They should market (and engineer) strongly on privacy since that's where Google is weak.
How can you build a privacy oriented search engine and still make money?
Had to look that up. I found https://help.duckduckgo.com/results/sources/ Bing is just one of "hundreds of vertical sources delivering" results to DuckDuckGo.
[1] https://arxiv.org/PS_cache/arxiv/pdf/1107/1107.5728v2.pdf
I meant that if Microsoft launched a privacy-focused product and hid their involvement with it they would receive extremely negative publicity on the sort of websites people who use privacy-focused products read.
Completely leaving aside the harm that would cause to the Microsoft brand, it would also be completely useless. Essentially no one would switch from DDG to Microsoft's (hypothetical) shady clone.
Your hotel examples aren't relevant because people searching for hotels and people searching for private search engines don't evaluate details about corporate ownership the same.
We've gotten so far into unlikely hypotheticals I don't find this conversation interesting any more. I won't reply further in this chain. Have a nice day!
> To do that, we've developed an open source Instant Answer platform called DuckDuckHack,
which links to https://duckduckhack.com/ which says "DuckDuckHack is now in Maintenance Mode".
And the "four hundred sources" link links to 400 special case replies. They are probably useful, but fire rarely. It's basically Bing, and that page is a bunch of spin.
I have always thought DuckDuck were doing their own searches.
It might be worthwhile to do the search part in-house and outsource the question-answering functions. Wolfram Alpha and IBM Watson could be used for answering common questions.
https://anolysis.privacy.cliqz.com/
As you can see in this subthread, they claim it's "anonymized" analytics,
https://news.ycombinator.com/item?id=21718694
which, as 99% of research suggests, isn't anonymous at all:
https://www.fastcompany.com/90278465/sorry-your-data-can-sti...
Hi, I read the thread and thought the answer was good enough, but it seems that you are not yet convinced. Let me try:
1) Here there is a list of publications regarding privacy by Cliqz (including published scientific papers). It should have fairly easy to find it using a search engine :-) https://0x65.dev/pages/dissemination-cliqz.html Hopefully, the paper will convince you that Cliqz privacy commitment is serious.
2) Feel free to monitor your own traffic to see whether or not we are tracking you.
3) Honestly, if someone tells you that anolysis means anonymous + analysis, why do you not believe it? It does not take long to find references of the name on the source code. On a separate note, as a company (Cliqz) that offers anti-tracking and ad-blocking, I can tell you that blocklists are a bit more sophisticated than that.
Hope that this will address your concerns,
[comment edited: why do you not believe it?]
For me it was the data proxying through FoxyProxy that made me uncomfortable. I have also remained unconvinced about the motivation for not using Tor: because it is hard to integrate into extensions. Cliqz also has its own browser, not just an extension, where they could opt to use Tor, but they route through FoxyProxy.
You must find a way to make it technically impossible to identify users, a legal or business structure is not enough. It wouldn't be unheard of to secretly own the proxy company and relink user data.
(EDIT/Disclaimer: I obviously work for Cliqz)
The openness must be paired with privacy that is guaranteed under all circumstances, and the most common way to achieve that would be to route the anonymized data through Tor.
Your search engine being also available on Tor has nothing to do with data collection by Cliqz on other sites. Your search engine website is not the avenue through which the Cliqz browser and extension collects data as you browse the web, I'm confused why would you even bring it up.
Disclaimer: I work for Cliqz.
Would be really interesting to know your concerns with FoxyProxy.FoxyProxy is legally bound to not log the IP or share it.
From the HPN protocol's perspective, data can be routed via any trusted party - in our case it's FoxyProxy.
Right now, there is no way to configure this in the Browser, but should be doable. It's actually one of the motivations to move to the newer version of HPN[1].
We do agree that sending data through Tor network is the gold standard for anonymity.
- We did a lot of work on getting Tor running in Cliqz browsers. It's a hard problem but definitely do-able, something we might pursue again in future[2]. - We also have experimented with WebAssembly version of Tor client to make it compatible for web extension[3].
Having the ability to use the Tor network in Cliqz products is also good, because we can actually leverage the anonymity guarantees by sending data via .onion services. You can also check more details under evaluation section of the paper[4].
In case you wish to check the network traffic you can also check the debugging section[5].
References:
1. https://www.0x65.dev/blog/2019-12-04/human-web-proxy-network... 2. https://github.com/cliqz-oss/browser-f/commit/12fcef8479d9c3... 3. https://github.com/cliqz-oss/browser-f/commit/12fcef8479d9c3... 4. https://arxiv.org/pdf/1812.07927.pdf 5. https://www.0x65.dev/blog/2019-12-03/human-web-collecting-da...
It does not take long to find references of the name on the source code.
No references in any of your published source code, and your search engine isn't free software:
https://github.com/search?q=org%3Acliqz-oss+anolysis
Honestly, if someone tells you that anolysis means anonymous + analysis, why do you not believe it?
Cliqz has done unsavory things in the past (like the Firefox fiasco a few years back, for example, which I can't fault Cliqz entirely for: Mozilla is just as guilty).
On a separate note, as a company (Cliqz) that offers anti-tracking and ad-blocking, I can tell you that blocklists are a bit more sophisticated than that.
"anolysis" gets around both uBlock Origin and uMatrix, despite both of them automatically blacklisting any URL with "analytics" in it, as an example. Getting around the most popular content filterers on the internet is a pretty strong signal.
You looked at the wrong tab, check in "Code" to find "code" related to Anolysis: https://github.com/search?q=org%3Acliqz-oss+anolysis&type=Co...
> Cliqz has done unsavory things in the past (like the Firefox fiasco a few years back, for example, which I can't fault Cliqz entirely for: Mozilla is just as guilty).
Not sure how this is Cliqz' fuck-up. We are not hiding anything. On the contrary we are very transparent and detailed about how everything we do is designed to not track users. All of this is on our new tech blog: https://0x65.dev, feel free to have a look between two comments on HN and give us some feedback!
> "anolysis" gets around both uBlock Origin and uMatrix, despite both of them automatically blacklisting any URL with "analytics" in it, as an example. Getting around the most popular content filterers on the internet is a pretty strong signal.
It's not called "getting around it" when there is no tracking or ads going on (if you want to see how smart the "most popular content filterers are" check out this link and see that the image is blocked because it contains the substring: "analytics": https://whotracks.me/blog/private_analytics.html. Wicked smart!).
Anolysis is not a typo, it's a project name, people tend to do that when they care and spend a lot of time on projects: give them names. So, at the risk of repeating myself, Anolysis = Analysis + Anonymous (at the time we thought it was a pretty neat name!).
Anolysis does not operate outside of Cliqz products (no websites analytics here and we do not rely on a third-party, we built it in-house for this reason) and we put a lot of work into it to make sure it does not use a unique ID (like virtually every other analytics out there) but allows to by-design not track any single user (in fact the system does not even have the concept of a user). Sure, we did not write extensively about it but I guess we have to start somewhere (in December we are writing on 24 different things we do, we will be sure to consider Anolysis as a good candidate for a technical blog post in the future).
What you attribute to malice is simply a lack of time, as you probably noticed Cliqz is working on solving a lot of very hard problems (search, browsers, antitracking, adblocking, privacy-preserving telemetry and so much more) and writing a paper about the new system you designed and implemented is not always the priority :)
You're also owned by a media company, that makes it even harder to believe that you're going to respect users privacy.
Add to that the tone of arrogance of articles such as "the world needs Cliqz", you can see why it's a no-no.
Every few days there is a post on HN trying so hard to convince the readers that Cliqz is the best, even though the articles read between the lines, that the Cliqz team does not have the capability to make its own search algo or make slightly more complicated queries.
I am experienced enough to know where this is coming from: managers that do not know what they are doing and engineers drunk from glory that do not see their own mistakes.
Please Cliqz hire a search engine expert. Hire great engineers, they are going to cost twice, but you're going to get a search engine that actually works.
Please, or the HN community will have to bash you every time you post an article.
They’re basically reverse engineering Google by looking at user logs.
Google will always have a leg up here because they have all the Google data.
And even if it does work for a while, there still needs to be the original signal to copy. Someone will have to crawl the web and index content.
I’m super eager to find new approaches to search, but another Google clone is not that.
A sufficiently strict intellectual-property regime might find this a copyright violation. But without any inside info, I strongly suspect Google & other incumbents already do similarly-indirect modeling of their competitors' behavior, via extensive query/click-trail mining, in ways that ultimately feed into improvements of their own systems. So, they might not want to press the issue.
Still, this creates a dependency on the competitor you were hoping to displace, where most of your earlier values comes from "drafting" in the easy-path they've already cleared.
Copying, learning would be a bit more precise, and I'm not kidding. We do not answer queries 1 to 1, query-logs are used to build a more concise model of the page. It's not just a cache, we started like this, but we quickly learn to answer unseen queries.
A sufficient strict IP might find text snipped a copyright violation too. It's a trick area. Personally, I'm at peace as we get the content of web pages, as everyone else.
As for the dependency, it would be if we were not able to generate our own synthetic queries, which we are. So even if all other search engines of the world were to disappear, we would still be able to operate. That was not always the case, as you pointed out.
Interesting you bring this up. Wasn't the company funding cliqz in favor of text snippets being copyright violations when the big G does it? (German/EU Leistungsschutzrecht) [0]
I mean, I think bootstrapping from google query logs is fine, but following the money it does seem like a double standard. Cliqz "stealing" from Google SERPs is fine, but Google news "stealing" text snippets isn't?
[0]: https://leistungsschutzrecht.info/stimmen-zum-lsr/pressearti...
(Relevant quote: Das Leistungsschutzrecht halte man nach wie vor nicht für falsch. Man setze sogar weiterhin auf ein euop. Leistungsschutzrecht, mit dessen Hilfe man sich erhofft, endlich Geld von Google zu erhalten.
sloppy translation: We still consider the [we-want-money-for-google-news-snippets-law] to not be the wrong approach. We in fact contiue hoping it will work out on a european scale)
As all German media? Yes, it seems that they were lobbying for it, not clear if they still are part of the consortium or not. Why? No idea, really. Perhaps they are diversifying, lobbying on one hand, trying to build a competitor with another. But to make an assessment of the "goodness" of their intentions I prefer to stick to the facts: besides complaining to regulators, or not, they do fund a potential alternative. That's very commendable, cannot name many companies that are crazy/adventurous enough to put a ton of money on something as risky as what Cliqz is trying to do. To sum up, if the lobbying is a minus, the building is a massive plus, a clear positive outcome IMHO.
If that’s not a Google clone I don’t know what is.
They're looking at how users use Google Search because the data's there. They're making a competitor to Google Search. That doesn't mean they're rebuilding Google Search's SERPs, or making a Google Search “clone”; I've got results from Cliqz for queries I'm confident have never been put into Google before, meaning it's functioning as an independent search engine.
This. Having worked in past life for one of their competitors, can confirm - what users click on (in SERP) is one of the most powerful signals for ranking. And who got (almost) all the clicks in the world? Google!
That's why it's so damn hard to beat them. It's the unreasonable effectiveness of data: more data (which they have almost all of) usually beats a smarter algorithm, and with 20 years R&D, theirs is surely not dumb.
Do the clicks belong to users or Google, that's an interesting question, though.
So there are Google and Bing - any others?
Qwant is a France-based search engine that has a similar privacy focus to DuckDuckGo.
Million Short tries to find sites that other search engines miss.
Wikipedia has a more comprehensive list:
In 2013 they said they were temporarily using Bing results they purchased as "training data" [wikipedia].
They haven't given any updates on that I could find (in English, at least), but they recently proudly announced they have 20B pages in their index in an article on their partnerships with Microsoft [betterweb]. Google's index is 1500 times larger, so I'm not sure how competitive their own index is [goog]. And, if they no longer needed to rely on Bing results, wouldn't they announce that?
[wiki]: https://en.m.wikipedia.org/wiki/Qwant
[betterweb]: https://betterweb.qwant.com/how-microsoft-tools-strengthen-q...
[goog]: https://venturebeat.com/2013/03/01/how-google-searches-30-tr...
I'd say they need to start somewhere. Using other search engine's results is a ways to get things started, so that they can build their own index on crawled content later.
For now, it's great to see another competitor for Google coming up.
It probably can't happen now, since there are billions of websites, but it was a simpler time, and finding something you needed wasn't THAT hard.
The best idea anyone had, which I thought wouldn't scale but didn't have better ideas, was implementing a keyword registry à la AOL.
Now compare the world of search to the world of mobile phones. Imagine if mobile phones only came with proprietary apps, and there were no app stores. That's where search is right now.
Further, they don't have to be better than google in the quality of their search results. As soon as the results have good relevance, it's good competition.
What matters is that they are 'good enough', that there is a hint of competition to the Google-Bing monopoly. Offering privacy centric 'competition' is what they say their main aim is and what their success should be judged upon.
> A web search engine or Internet search engine is a software system that is designed to carry out web search (Internet search), which means to search the World Wide Web in a systematic way for particular information specified in a textual web search query.
Just because they don’t do the crawling like Google and Bing doesn’t mean they aren’t a search engine.
Now any service with a name like that seems shit compared to more readable/writeable names: Stackoverflow, Quora, datadog.
Case in point is srht.co, which rebranded to Sourcehut for similar reasons.
* gugal -- like frugal
* gugle -- like bugle
* ghougle -- like ghoul
* googull -- like gull
[1]: https://whotracks.me/blog/how_cliqz_antitracking_protects_us...
[2]: https://static.cliqz.com/wp-content/uploads/2016/07/Cliqz-St...
You might as well name yourself "Nigerian Princes Inc."
Enjoy 1 nut.
( Extreme comment fyi )
Maybe I'm old school, but I don't want software that fixes my spelling mistakes. I want software that fails when I make a mistake.
I feel like the web parsing / indexing, perhaps rather than the search algorithm itself, is the hardest part of rolling a new search engine (largely due to the associated computing and storage costs).
Will probably keep an eye on this blog.
In our case the "queries" are also the index creation components. Every time someone discusses something, we are indexing it, so you can search media, documents, people from context. We hint at how this works here: https://austingwalters.com/fast-full-text-search-in-postgres...
The downside of our approach is it needs lots of conversation data. From their TLDR version:
"""
- Our model of a web page is based on queries only. These queries could either be observed in the query logs or could be synthetic, i.e. we generate them. In other words, during the recall phase, we do not try to match query words directly with the content of the page. This is a crucial differentiating factor – it is the reason we are able to build a search engine with dramatically less resources in comparison to our competitors.
- Given a query, we first look for similar queries using a multitude of keyword and word vector based matching techniques.
- We pick the most similar queries and fetch the pages associated with them.
- At this point, we start considering the content of the page. We utilize it for feature extraction during ranking, filtering and dynamic snippet generation.
"""
It appears 0x65 has similarly figured this out, the name of the game is forming proper search queries. In their case, their results would be good as soon as they start indexing and create synthetic queries. IMO might be better for documents and what not.
Either way, interesting to compare notes! Kudos to the work.
I'm trying to implement a faceted search in postgres and currently using window functions to count subcategories (a la http://akorotkov.github.io/blog/2016/06/17/faceted-search/), but not sure if it's the most efficient.
"SELECT planrows FROM estimate_row('SELECT COUNT(*) ON table WHERE XXX')"
If you're ever looking for something to write about for a new blog post, I would love to learn more about how you implemented that estimate_count function.
Thanks for the tip in the right direction tho!
They're doing an "Advent Calendar" series where they're posting one a day:
https://news.ycombinator.com/item?id=21676252
https://news.ycombinator.com/item?id=21684708
https://news.ycombinator.com/item?id=21694980
and https://news.ycombinator.com/item?id=21716860 (not even a day ago).
This is a problem for HN because users here are not used to this sort of repetition—indeed, we moderate HN explicitly to dampen repetition, because the point of the site is curiosity and curiosity withers under it (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...). The more these posts show up with what for HN is a crazy frequency, the more likely users here are to experience it as a barrage and start to complain.
On the other hand, these articles are well-crafted, contain a lot of information, and would normally be fine HN submissions. The topic of building a new search engine is intrinsically interesting. It also resonates with a lot of themes that get discussed a lot on HN (concerns about big tech and so on). So this is a different situation than the usual marketing onslaughts that HN gets subjected to, where the content is crappy, users flag it away, and moderators squash what users missed.
I'm not sure what to do about this yet.
>> "The total size of our index currently is around 50 TB."
Could you share your current index size (number of pages, size of raw text) to put those 50 TB into perspective in order to get an idea how much less resources in comparison to your competitors you need? This would help to compare your approach to Elasticsearch, Solr, Lucene
Cliqz bought Ghostery to acquire a pool of privacy-conscious users. The goal is to show them ads. Not sure how excited they will be about that.
If Cliqz really is a search engine, can a user submit a query to the database using her own choice of tcp/http client. It looks like submitting requires first downloading and installing software from Cliqz.
"It may seem like Common Crawl would suffice for this purpose, but it has poor coverage outside of the US and its update frequency is not realistic for use in a search engine."
The 10PB of disk is also quite reachable given its possible to buy bulk 10T disks at $150 each.
Bottom line, I did some of these calculations a couple years ago because I was interested in a topic based search engine that only indexed for certain topics and basically tossed any crawler results that didn't appear to fit the subject matter.
So, while the web is a lot bigger than when google started, storage and compute is also a lot cheaper. A web search engine that specialized in say cooking recipes might be entirely doable on a fairly limited budget.
But...my searching has gotten to the point where well over half the time I am no longer looking at Google's conventional search results.
Thank you for bringing this up. Although, this is not relevant in the context of the blog post, we are on that list by mistake. We do NOT collect any personal data in our browser: (more details e.g., here: https://0x65.dev/blog/2019-12-02/is-data-collection-evil.htm... and https://0x65.dev/blog/2019-12-03/human-web-collecting-data-i...) and we go a long way to make sure not even implicit indentifiers go through. We believe we ended up on that list for a bad Firefox experiment and we will reach out to the maintainers, make our case.
Disclaimer: I work for Cliqz.
so make a competitive "suggested autocorrect" solution and then I think you'd have a stew going.
[Disclaimer: I work at Cliqz]
Although it is anonymous data - currently we are not aware of any de-anonymization attacks - it is still data that came from real persons. We have a responsibility: once the data is out, we have to guarantee that no-one will ever be able to identity a single person in the data. Take also in account that attackers can combine multiple data sets (Background Knowledge Attacks); that even includes data sets that will be published (or leaked) in the future.
You should never be too confident when it comes to security, neither should you underestimate the creativity of attackers. What we can do - and did in the past - is to simulate the scenario in a controlled environment by hiring pen testing companies. If they would find an attack, they will not use that knowledge to harm the persons behind the identities that they could reveal.
That is the main reason. We don't want to end up in a situation as AOL or Netflix when they published their data. By the way, Netflix is an example of a background attack where they needed to combine data sources.
There is also another argument. Skeptics will most likely remain skeptics, as we cannot proof that we did not filter out data before publishing. In other words, there is nothing to gain for us, we can only loose. Trust is important, but for building trust, it is better to be transparent about the data that gets sent on the client. You can verify that part yourself and do not have to rely on trust alone. That is the core idea behind our privacy by design approach.
Those are the arguments that I'm aware of why we will not open the data. However, getting access in controlled environments is possible. If you doing security/privacy research, you can reach out to us. In my opinion, having more people that will try to find flaws in our heuristics is useful. That gives us a chance to fix it before it can be used for attacks.
One notable exception: https://whotracks.me is built from Human Web and all its underlying data can be freely downloaded. We know that it has been already used for research.