285 karma · joined June 7, 2012
I hate to answer this one, becasue it looks too much marketing-speech but this feature exists. Not on beta.cliqz.com but on the drop-down search on Cliqz browser.
Based on the tabs you have opened, different query expansions are selected. For instance as you type "hotel in ma..." probably would show you results for Mallorca, but if you have "Madrid" on a tab, then it will show results for "hotel in madrid".
There will be a blog post about this contextual search because it's our showcase that is possible to do personalization without compromising privacy. All this is done privately, the browser receives results for multiple expansions and can chose which one to display based on local information. We never track or collect sessions of users.
FYI, there will be a bunch of articles regarding search in the next week.
Your point is spot on. Old pages tend to have more association to seen queries, which does not play in favor for new pages.
That said, however, there are a couple of things to consider: 1) seen queries is not the only way to create queries, we are pretty good creating synthetic queries based on the content, descriptions, etc. This queries are more noisy that the seen queries of course, but good enough. And 2)novelty, freshness and popularity are very important features on the ranking. Feel free to try out any new topic you might think of on https://beta.cliqz.com, you will see that is not only "stale" content.
The salaries figures, 25K and 20K seem extremely low to me. Truth be told, I've been abroad for the last 6 years, but back then, no way we could hire anyone decent for less than 35K (Barcelona).
I would advise that you are going the earn 20K/25K, do not try to go unless you really want to survive like a student :-) Do not see how are you going to pay the rent with that salary, check prices online.
Both cities are very nice and you will enjoy them both.
Last time I lived there, Barcelona was more start-upy than Madrid, and Madrid more corporate than Barcelona. In any case, they both have plenty of tech, I would not worry about that. You will enjoy your time in either city, if you earn enough of course. Make sure of that first.
But please focus on Google, which is the one putting the money!
Why Google pays so much money if they already have the best product and what the users want? Are they just giving B10$ as charity?
There are many reasons for doing that, but none of them is aligned with the make-believe statement that "build a better product and people will come"
I would agree that is a necessary condition, but by no means sufficient, which is kind of sad.
Yes, we are not the only approach. Even when we started collecting data back in 2014 it wasn't the only one. But we did not find any suitable off-the-shelf solution back then.
Homomorphic encryption was discarded because some data to be send needs to be on the clear. For instance a url we need to fetch, so cannot transform it. Also, computationally is very expensive.
Federated learning is actually closer to what we do, take our approach as federated learning where each node is a single user and where all aggregation of records need to happen there, so that record-linkage on the final collector is impossible.
Tomorrow and the day after tomorrow we are releasing the technical details on different blog-posts, this one was more a motivation/introduction to the main dish.
I'm sure it's obvious that I work at Cliqz and on this very topic :-) Tomorrow there is a more technical article about our data collection, hope you like it.
Also, I would like to add that any methodology applied to protect the data of the user is welcome, does not have to be ours at all. There is one caveat though, the privacy protection has to be on origin, a solution that send data that then has to be anonymized is no good in our book; because there is no guarantee that the raw data is removed.
For our use-case we believe it was the easier way to get the data we needed while respecting the privacy of the user.
Comparisons of quality are relative, Google was good from day one because it was better than the competitors at that moment of time. Problem is that if one starts using X, and that provides an edge, the other have to follow just to keep up. A very unfortunate rabbit hole.
I'm with you 100% that personal data makes search worse in in many cases, too much personalization is detrimental.
One last thing, which I'd like to stress, is that data from users is not necessarily the same as personal data. For instance, you can use the data from users as if they were just sensors, without trying to build sessions out of it.
In a way it's similar to what Google made when they started. Google was the first and only to use anchor text, which is user generated and a less noisier description of the content of the page. (True that they did not use the users themselves but they did the proper automatic crawling). But in a way they collected sensor data from users, not personal data. That's what we should still be doing today. The problem is , however, however, is that is too easy to go beyond "sensor data" and start to collect full sessions and even personal data.
The other option is that Google does not pay to get the Apple users, which might come otherwise, but to prevent Apple to try something funny, either directly or indirectly.
In any case, the "build a better product and people will come" mantra is flawed when companies in the space pay each other billions for distribution.
"Google's dominance is almost entirely due to the fact that its by far the best at search."
is, and I'm sorry to say, a make-believe
Google's quality is better than anyone else, that is a fact. Let's go to the point. [[I work on Cliqz, and in search]]
Do you know how much Google pays apple to be the default search engine?
According to you, nothing, because people will go to Google because it's the best.
Well, it turns out that last year was more than 9 Billion (with a B). Quite a lot of money poorly spend, someone should really get fired :-)
We are all ears on ideas on how to beat Google :-)
But on a serious note; i do not think that the aim is to beat Google because of Google but because on the monopoly they have on the access to the information. And that's why we are building a potential alternative there.
Attacking from a different domain, might beat Google on that area, or let's go wild here, perhaps even on revenue, but would not fix the problem we wanted to fix to the begin with.
"Replicating" 99% of the functionality with a quality to be good enough is a very valid option IMHO. Otherwise if one decides to leave Google, where do they go? If they do not want to advertise on Google, where do they go? There are too few alternatives, there are many names, but most of them are aggregators (full or partial). That said, if besides Bing, there would be 4 or 5 like them (or better), then yes, I concede that Cliqz as is, would make no sense.
Whether Cliqz would become a bully like Google if successful or not, I believe it's impossible to answer.
But at least, we know, that the data we collect from our users to build the search engine is totally anonymous and is used only to build search and the browser.
So, in the case that Cliqz would turn evil, I would stop using it, knowing for a fact that there are no sessions about me on their data, none.
Someone on the thread mentioned fiduciary duties, true that companies must maximize profit, but that does not mean becoming a gear on global surveillance conglomerate.
However, you have a point that if you are to use this setup, you could be better off using mongodb right away, it's a matter of trade-offs.
A use-case of JOR is, for instance, when you need to ship a system on the premises of your customer, or you have an embedded system. In such cases, the ease of configuration and maintenance of redis really pay off (we run both systems in production). You do not have all the power of MongoDB, only a fraction of it, but you also have a fraction of operations cost.
But I do agree, that if your system is big enough, or you have the resources for the proper mongodb setup, go for the real thing (mongodb) :-) JOR is some sort of middle approach, for those cases that you say, would be good to have the queries from mongo but it's just too much hassle to install it (in-house or in-premise).
It seems that we have found this unusual conditions you mention :-) Before compression our replication was often lagging (when below the required bandwidth), which is very dangerous, even if after a minute bandwidth goes well above the threshold.
The case of replication is kind of worst case, it's no good to have an average of 30 Mbps if it stays a substantial fraction of the time below the threshold. For every hours at 75% memory grew 2GB, but even one minute below the threshold is bad... if the replication queue builds up and there is a crash, consistency goes out of the window. So better to over-dimension capacity.
We even did a quick test with Hetzner instead of Rackspace, it was worse. The average over 4 hours was about 15 Mbps.
However we also use ssh tunnels for our continuous integration (jenkins) and the irc bots that we use for development (http://3scale.github.com/2012/06/29/irc-driven-development-p...).
In this setup we have the autossh going awol every one or two weeks. But as the blog post mentions, the ssh tunnel runs between our HQ (fiber) and Amazon, not very reliable. Furthermore, there is a lot of idle periods, which seems to trigger most of the issues. We cannot give more specifics since we forcefully just restart the daemon with a monit/munin combo.
however, getting out of cloud does not help. The problem is the high-availability. To maintain 99.9% we cannot rely on a single data-center, no matter what the claims on availability zones say. We have even seen network partition which are the worst of the worst. Once you go on the "Internet" there is no guarantees (unless dedicated lines).
I suspect (guess) that not having an idle connection help the tunnel to not randomly drop.
Initially we didn't add compression to save bandwidth but to solve the network bandwidth issues between amazon and rackspace. However, it turns out that you can save on the bill quite a lot, in our case about ~ $700.
Regarding the binary, would be very nice to have. The first though is that compression would work better for us than binary since our keys are quite long (50/100 bytes) and very regular on their patterns, so compression works very well in our case. Probably can be extrapolated for other cases where the keys are typically heavier than the values.