HNHacker News
TopNewBestAskShowJobs

tschakkaMarc

19 karma · joined November 20, 2014

Lives, defends and takes feedback for @cliqz Reachable at Twitter @MarcAlHames
submissionscomments
tschakkaMarc··on Building a Search Engine from Scratch
Or via Tor browser as Onion service: search4tor7txuze.onion (quite unique - if available it will show onion services instead of www domains).
tschakkaMarc··on Building a Search Engine from Scratch
Marc from Cliqz here: it’s something I’m thinking about a lot since years, it’s a very clean model. We even did some tests, but unfortunately the majority of people is not like hacker news. What is a reasonable price point for you per month? How would you pay (ie would you feel comfortable giving us your credit card and create an account? (I don’t think crypto is big enough yet to solve this)
tschakkaMarc··on Building a Search Engine from Scratch
Well there are people who believe this and others who believe that ... as I said - very strong opinions (like yours) on both sides of the aisle even internally. Interesting proposal with the Nigerian Inc. - if we ever launch an email service or - better - a Spam filter, I’ll make sure to name the company Nigerian Princes Inc.!
tschakkaMarc··on Building a Search Engine from Scratch
@MarcAlHames on Twitter
tschakkaMarc··on Building a Search Engine from Scratch
It’s the one time where I can honestly say: The discussion we’re having within Cliqz about our brand name are even more heated and controversial than here on Hacker News (and Reddit for that matter) ... but then there is this saying: “Every brand name is shit until you surpass one billion users - than it becomes brilliant”. More seriously: we do think about it a lot - happy to get ideas.
tschakkaMarc··on Building a Search Engine from Scratch
Well, we released our beta search for Tor yesterday: search4tor7txuze.onion/ (works obviously only in the Tor browser). That’s as good as it gets regarding making it technically impossible, isn’t it? More complex obviously for a browser - but the answer can simply not be „only no data at all is good“, because that’s a destructive approach that only favors the worst privacy intruders; no one would then be able to build up competition. We work hard to be as transparent as possible about what we do. Show me any other company that builds a big data product (like search) that is so transparent about „yes we collect data, but it’s non personal - here is how we do it, please scrutinize us“. Are we perfect? Hell no! But we try and we go a long way to be challenged to improve. If we were shady - would we be naked in front of you, showing each step we take? We would not even interact with the tech folks (especially not on hacker news, where people really know what they talk about), but scam people who know less (at least that sounds like a more reasonable strategy to me if I would want to fool people - which we don’t).

(EDIT/Disclaimer: I obviously work for Cliqz)

tschakkaMarc··on Show HN: New, Fast, Anonymous, Tor-First Search Engine
Thanks for the comment. Yes - it will only work in Tor or in a browser with a Tor tab (e.g., Brave).
tschakkaMarc··on Human Web – Data Collection Without Privacy Side-Effects
Hi – thanks for the feedback. I’m Marc (disclaimer, I work at Cliqz). The goal of the collaboration was to jointly build a better, more private search engine. Don’t forget that every major browser today sends every keystroke of the Omnibar to either Google or Bing. No privacy mechanism in place (and don’t even get me started on all the tracker madness). We wanted to jointly replace that. Cliqz is building a new search from ground up with privacy by design. In the end the collaboration didn’t work out (for many reasons, lack of privacy was not one of them though). Firefox changed back to Google as search provider. Coming back to your point: There are many services that can’t be built without data. Search is one of them, without data you will have a very bad search engine, impossible to compete. We explain this in detail here: https://0x65.dev/blog/2019-12-02/is-data-collection-evil.htm... . We took maximum scrutiny, and this article about Human Web is exactly there to explain how we collect data that is needed, without the side effect of collecting personal data. We are so transparent about this, because we want the scrutiny. Our business does not depend on collecting personal data or actually any data. But our product needs a lot of data. Denying anyone to collect data – even if they are as open, transparent, and without any interest in personal information – just means you support those that are the incumbents and have no interest in privacy.
tschakkaMarc··on The world needs more search engines
And here are all details how we remove all personal information with a technology we call Human Web: https://0x65.dev/blog/2019-12-03/human-web-collecting-data-i...

We love to get scrutinized and get feedback on this - we’re very serious about privacy.

tschakkaMarc··on The world needs more search engines
There were extensive tests with opt-in before (Testpilot), but these are super biased towards techies/enthusiasts (by definition if you read HN or use Testpilot you’re not representative ...). At some point you need to both test and get data from more mass market and that would never work with opt-in. Hence the scrutiny about not even technically be able to do record linkage etc. And some of the measures you mentioned are/were applied (we post about this in the next days).

I also stick to my original point: those users who had cliqz had significantly more privacy than those without.

Having said that: I don’t think, you and me are that far away from each other. But: If we, who care about privacy constantly criticize or even shout at those who also care about privacy, those who build better products, but maybe don’t follow an idealistic “no data at all paradigm”, then we will always end with the worst data collectors, because non of the alternatives will ever have a chance (or people get frustrated and decide they can make more money at Google or ad tech).

By the way, we have a post about data and how we collect it in our blog today: https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil... - you might find it interesting).

In any case thanks for challenging us. I don’t believe we’re perfect. But we’re trying!

tschakkaMarc··on The world needs more search engines
One doesn't exclude the other. I admire a lot of what Google has done and agree on a very strong execution. This still doesn't change the fact, that a search engine (1) gets better with every query-click pair and hence favors the market leader, (2) is a highly profitable business, which allows to buy distribution (this is the biggest market entry barrier), and (3) there are many small points that add up (crawler access is one minor but annoying issue: Web-Sites very often see all crawlers except Google and Bing as "bad scrapers"; so small companies need to invest a lot of time to convince them to get access and many will still never allow it what users then rate as bad quality) ... but, yes, Google is doing a lot of things right, they have very good people. Many of them my friends or ex-colleagues. This still doesn't change the fact of the original article: A 93% market share is not great.
tschakkaMarc··on The world needs more search engines
Yes, I’m defending it, because again: We took drastic steps to never send anything private (like checking within the browser whether the URL is unique or different if logged in or out and then never sending it, not to mention that there of course was no identifier and we made record linkage impossible, so no click profile, and much more). If in doubt we drop and don’t send. And again – there were tons of (pen) tests and scrutiny to make sure no private data point ever leaves your browser. It is built with the mindset “if it reaches our server, we should technically not be able to identify any single person or any surf pattern or any private URL” – this was and is also tested by many (privacy) researchers before and after the experiment. And again, all this was and is open source. This is way more than any industry standard, and I simply don’t know of any company that works with data that has a higher standard. Be our guest to validate it yourself. And please read our blog post Tuesday: we will explain how this is done. But if you simply oppose this (and similar methods from people who really care about privacy), you basically accept the status-quo of the worst data collectors, because no one else then will ever emerge (because you do need this kind of data to build a search).

[EDIT]: Just to clarify and not have anyone create the wrong idea - I defend my earlier point. But your question is loaded. Here's why: We do not collect browser history, which by definition implies being able to piece visited urls back to a profile in our servers. That is impossible - to us, each single URL comes as a detached datapoint - devoid of any information that can be used to aggregate them back to a user profile.

tschakkaMarc··on The world needs more search engines
Hi, Marc from Cliqz. This one haunts us (in HN and also in my dreams). Let's first get one thing clear: It was a terrible blog post. Second: The 1% who had Cliqz installed were actually safer than the ones without. Why? If you use a (most) browsers every keystroke in the URL bar gets send to Google. This is how autosuggest works. This was and is the case with Firefox also. Cliqz in the functionally implemented in Firefox was not that different. Except - it comes with privacy by design: We actually proxied those requests to not get the IP, we take special measures to not collect private data in the first place. And all this was tested and scrutinized a lot before and after the test. And if you don't trust it or believe it, it also happened to be open source. In the end, you were better off with Cliqz than without. Third (and maybe most importantly): If you don't support anyone who "collects" data, even if they do it in the most transparent way, without collecting any private information, in fact going a long way to delete all PII, open-source and with privacy by design, someone who has no business model built on collecting profiles, then you only criticize Google, but will never have an alternative or only alternatives that white label those that do collect all data. Because building a search without data is impossible. We were hoping that together with Firefox we would build a better solution. For some reasons this didn't work out, but lack of privacy was never one of them (lack of a business model more so, because sending every key press to some is more profitable than others, regardless of how private each is). In fact, our way of doing things is much more private than most people out there who claim to be private. This includes most browsers and search engines. And all of this is open source, so verifiable. And while we speak about it: In our blog https://0x65.dev we will over the next 24 days publish pieces on what we do and why we do it and why it is important, data collection and anonymization will be a big part of this. Last not least: All this is clearly not your fault, so my rant might be inappropriate, but as I said, this thing haunts us for the wrong reasons. (Me and the team are happy to answer any questions).
tschakkaMarc··on The world needs more search engines
Hi, Marc from Cliqz (and one of the authors of the linked article): We at Cliqz do try and others (like mojeek in the UK or Seznam in Czech) are trying as well. But yes, it has very high costs. We did a lot of (I find smart) technical shortcuts (and will explain how in the next days in our blog), but it is still an investment. The main hurdle is however distribution. As said in the article, Google spends 1/4 of their revenue to block market access and this makes the investment case nearly impossible. It's their strategy to stop competitive innovation. This no company can solve and why e.g., in Europe you do see the 93% market share of Google. This is where this becomes a political topic. In fact since search tends to be a natural monopoly and information is so important this is where politics has to regulate and ensure competition. In an ideal world there would be 4-5 competing search engines in every region ...
tschakkaMarc··on The world needs more search engines
Hi, Marc here, I work at Cliqz. We (unfortunately) don’t (at least not a reasonable amount). We have an advertising model, but it does not rely on any personal data. We also always think (and tested) a paid product, which would be the purest and (In my view) best form. Experience unfortunately suggests, that only very few people would pay. But, I would be happy to be convinced otherwise. Going back to the article though: This is exactly why we believe there should be a lot of competition and a variety of different alternatives. And one or the other might figure out a new approach towards financing search.