HNHacker News
TopNewBestAskShowJobs

solso

285 karma · joined June 7, 2012

CS at Brave (https://search.brave.com/) Twitter: @solso Homepage: http://josepmpujol.net/
submissionscomments
solso··on A New Search Engine
Yes, we are on some minor blocking lists because of our data collection, even though is anonymous (please check the articles about Human Web on https://0x65.dev/) sending data, no matter what data, is a sin that has to be punished. A disservice to you ask me, but what can we do. [Disclaimer: I work at Cliqz]
solso··on A New Search Engine
Not sure I get your point. But contextual search only works for the search within Cliqz browser, on the address bar dropdown, on the client space. The same approach cannot be done on the (web-page SERP page, beta.cliqz.com), because from a web we have no access to the tabs opened. It could only possible via tracking and user-profiling, which is something that we do not do, or want to do.
solso··on A New Search Engine
[Disclaimer: I work at Cliqz]

I hate to answer this one, becasue it looks too much marketing-speech but this feature exists. Not on beta.cliqz.com but on the drop-down search on Cliqz browser.

Based on the tabs you have opened, different query expansions are selected. For instance as you type "hotel in ma..." probably would show you results for Mallorca, but if you have "Madrid" on a tab, then it will show results for "hotel in madrid".

There will be a blog post about this contextual search because it's our showcase that is possible to do personalization without compromising privacy. All this is done privately, the browser receives results for multiple expansions and can chose which one to display based on local information. We never track or collect sessions of users.

solso··on A New Search Engine
[Disclaimer I work at Cliqz] There is a lot of systems under the hood, depending if it's the main index or the freshness index. But if I have to pick one as database it should be Keyvi (https://github.com/KeyviDev/keyvi).

FYI, there will be a bunch of articles regarding search in the next week.

solso··on A New Search Engine
[Disclaimer, I work at Cliqz] I cannot answer for the "true" motivation of the investors, but their pitch and actions so far are well align with the fight against monopolies narrative. Do they want to get return on investment (eventually)? I would assume so, and I believe it would be fair. I do not see them as mutually exclusive. Of course, this is my personal opinion.
solso··on A New Search Engine
[Disclaimer, I work at Cliqz]

Your point is spot on. Old pages tend to have more association to seen queries, which does not play in favor for new pages.

That said, however, there are a couple of things to consider: 1) seen queries is not the only way to create queries, we are pretty good creating synthetic queries based on the content, descriptions, etc. This queries are more noisy that the seen queries of course, but good enough. And 2)novelty, freshness and popularity are very important features on the ranking. Feel free to try out any new topic you might think of on https://beta.cliqz.com, you will see that is not only "stale" content.

solso··on A New Search Engine
An excerpt from the 1st post of this series: "Why would a team be motivated to build another search engine? Why would Hubert Burda Media finance this over several years (they continued to back us especially in times when things got tough)?" https://0x65.dev/blog/2019-12-01/the-world-needs-cliqz-the-w...
solso··on Ask HN: What are the Madrid and Barcelona tech scenes like?
Refusing to speak spanish is a myth. It's much more likely that you will not ever be able to learn catalan becasue poeple will always switch to spanish the moment they see you are a foreigner. Whether you want to learn spanish or catalan, both or none is up to you. You can live a normal live speaking english only in both places. Have plenty of friends that never got to learn the language beyond some basic pick-up lines and food-related sentences.

The salaries figures, 25K and 20K seem extremely low to me. Truth be told, I've been abroad for the last 6 years, but back then, no way we could hire anyone decent for less than 35K (Barcelona).

I would advise that you are going the earn 20K/25K, do not try to go unless you really want to survive like a student :-) Do not see how are you going to pay the rent with that salary, check prices online.

Both cities are very nice and you will enjoy them both.

Last time I lived there, Barcelona was more start-upy than Madrid, and Madrid more corporate than Barcelona. In any case, they both have plenty of tech, I would not worry about that. You will enjoy your time in either city, if you earn enough of course. Make sure of that first.

solso··on Show HN: Parrot.vc – I forced a bot to read 65,000 VC tweets and it became a VC
works well, it would have fooled me. The quotes are as bizarre as the real ones :-)
solso··on The world needs more search engines
Forget about Apple for a second, they will sell whatever to anyone, for sure.

But please focus on Google, which is the one putting the money!

Why Google pays so much money if they already have the best product and what the users want? Are they just giving B10$ as charity?

There are many reasons for doing that, but none of them is aligned with the make-believe statement that "build a better product and people will come"

I would agree that is a necessary condition, but by no means sufficient, which is kind of sad.

solso··on Privacy or Data, a Convenient False Dichotomy
[Disclaimer: I work at Cliqz]

Yes, we are not the only approach. Even when we started collecting data back in 2014 it wasn't the only one. But we did not find any suitable off-the-shelf solution back then.

Homomorphic encryption was discarded because some data to be send needs to be on the clear. For instance a url we need to fetch, so cannot transform it. Also, computationally is very expensive.

Federated learning is actually closer to what we do, take our approach as federated learning where each node is a single user and where all aggregation of records need to happen there, so that record-linkage on the final collector is impossible.

Tomorrow and the day after tomorrow we are releasing the technical details on different blog-posts, this one was more a motivation/introduction to the main dish.

solso··on Privacy or Data, a Convenient False Dichotomy
Of course client-side aggregation can still leak privacy, it needs to guarantee that there are no explicit or implicit elements that would allow record-linkage on the server-side. On the diff. privacy, note that you mention before you store the data, the problem is not only there, we want to prevent the data to be send at all. Diff privacy on the client-side is tricky if distributions are unknown. That's why we go for a "simpler" approach, all records send by any user should be unlikable, always. If aggregation is needed to satisfy the use-case, it can only be done on the client itself. Re-identifiability only becomes possible if a mistake is done, no a priori distributions are needed. IMHO this approach is easier than diff. privacy at the cost of being less expressive. Diff. privacy data allow for multiple use-cases where we do not (as all records) have to be independent from one another.

I'm sure it's obvious that I work at Cliqz and on this very topic :-) Tomorrow there is a more technical article about our data collection, hope you like it.

Also, I would like to add that any methodology applied to protect the data of the user is welcome, does not have to be ours at all. There is one caveat though, the privacy protection has to be on origin, a solution that send data that then has to be anonymized is no good in our book; because there is no guarantee that the raw data is removed.

For our use-case we believe it was the easier way to get the data we needed while respecting the privacy of the user.

solso··on Privacy or Data, a Convenient False Dichotomy
[Disclaimer, work at Cliqz]

Comparisons of quality are relative, Google was good from day one because it was better than the competitors at that moment of time. Problem is that if one starts using X, and that provides an edge, the other have to follow just to keep up. A very unfortunate rabbit hole.

I'm with you 100% that personal data makes search worse in in many cases, too much personalization is detrimental.

One last thing, which I'd like to stress, is that data from users is not necessarily the same as personal data. For instance, you can use the data from users as if they were just sensors, without trying to build sessions out of it.

In a way it's similar to what Google made when they started. Google was the first and only to use anchor text, which is user generated and a less noisier description of the content of the page. (True that they did not use the users themselves but they did the proper automatic crawling). But in a way they collected sensor data from users, not personal data. That's what we should still be doing today. The problem is , however, however, is that is too easy to go beyond "sensor data" and start to collect full sessions and even personal data.

solso··on The world needs more search engines
Of course it's business, but it invalidates the premise that people chose "mostly" because of the quality. If that was the main driver, Google could just pay zero, 9 Billion for few hundred million users is quite an hefty sum.

The other option is that Google does not pay to get the Apple users, which might come otherwise, but to prevent Apple to try something funny, either directly or indirectly.

In any case, the "build a better product and people will come" mantra is flawed when companies in the space pay each other billions for distribution.

solso··on The world needs more search engines
The claim:

"Google's dominance is almost entirely due to the fact that its by far the best at search."

is, and I'm sorry to say, a make-believe

Google's quality is better than anyone else, that is a fact. Let's go to the point. [[I work on Cliqz, and in search]]

Do you know how much Google pays apple to be the default search engine?

According to you, nothing, because people will go to Google because it's the best.

Well, it turns out that last year was more than 9 Billion (with a B). Quite a lot of money poorly spend, someone should really get fired :-)

solso··on The world needs more search engines
[disclaimer I work at Cliqz]

We are all ears on ideas on how to beat Google :-)

But on a serious note; i do not think that the aim is to beat Google because of Google but because on the monopoly they have on the access to the information. And that's why we are building a potential alternative there.

Attacking from a different domain, might beat Google on that area, or let's go wild here, perhaps even on revenue, but would not fix the problem we wanted to fix to the begin with.

"Replicating" 99% of the functionality with a quality to be good enough is a very valid option IMHO. Otherwise if one decides to leave Google, where do they go? If they do not want to advertise on Google, where do they go? There are too few alternatives, there are many names, but most of them are aggregators (full or partial). That said, if besides Bing, there would be 4 or 5 like them (or better), then yes, I concede that Cliqz as is, would make no sense.

solso··on The world needs more search engines
[[Disclaimer I work at Cliqz]] Your question is spot on, many of us ask ourselves the same question.

Whether Cliqz would become a bully like Google if successful or not, I believe it's impossible to answer.

But at least, we know, that the data we collect from our users to build the search engine is totally anonymous and is used only to build search and the browser.

So, in the case that Cliqz would turn evil, I would stop using it, knowing for a fact that there are no sessions about me on their data, none.

Someone on the thread mentioned fiduciary duties, true that companies must maximize profit, but that does not mean becoming a gear on global surveillance conglomerate.

solso··on A simple document datastore using Redis
Redis persistence is very reliable, reliable enough to be a proper datastore when using replication (master-slave). We have been using it in prod for 4 years, with many GB and not a single problem. From my experience, redis is perfect for storage.

However, you have a point that if you are to use this setup, you could be better off using mongodb right away, it's a matter of trade-offs.

A use-case of JOR is, for instance, when you need to ship a system on the premises of your customer, or you have an embedded system. In such cases, the ease of configuration and maintenance of redis really pay off (we run both systems in production). You do not have all the power of MongoDB, only a fraction of it, but you also have a fraction of operations cost.

But I do agree, that if your system is big enough, or you have the resources for the proper mongodb setup, go for the real thing (mongodb) :-) JOR is some sort of middle approach, for those cases that you say, would be good to have the queries from mongo but it's just too much hassle to install it (in-house or in-premise).

solso··on Having fun with Redis Replication between Amazon and Rackspace
The replication stream is about 20-25 Mbps for the current throughput of 25K-30K req/s, our peak to valley ratio is low.

It seems that we have found this unusual conditions you mention :-) Before compression our replication was often lagging (when below the required bandwidth), which is very dangerous, even if after a minute bandwidth goes well above the threshold.

The case of replication is kind of worst case, it's no good to have an average of 30 Mbps if it stays a substantial fraction of the time below the threshold. For every hours at 75% memory grew 2GB, but even one minute below the threshold is bad... if the replication queue builds up and there is a crash, consistency goes out of the window. So better to over-dimension capacity.

We even did a quick test with Hetzner instead of Rackspace, it was worse. The average over 4 hours was about 15 Mbps.

solso··on Having fun with Redis Replication between Amazon and Rackspace
that's totally correct. In the case of a redis replication, the ssh tunnel failing would be quite costly, since the whole database will have to be send again since the slave would do a SYNC op. And that can be multiple GB in our case. But so far, we haven't experience it.
solso··on Having fun with Redis Replication between Amazon and Rackspace
The glitches with the ssh tunnel were not for the redis replication setup. This has been running fine for 3 weeks already.

However we also use ssh tunnels for our continuous integration (jenkins) and the irc bots that we use for development (http://3scale.github.com/2012/06/29/irc-driven-development-p...).

In this setup we have the autossh going awol every one or two weeks. But as the blog post mentions, the ssh tunnel runs between our HQ (fiber) and Amazon, not very reliable. Furthermore, there is a lot of idle periods, which seems to trigger most of the issues. We cannot give more specifics since we forcefully just restart the daemon with a monit/munin combo.

solso··on Having fun with Redis Replication between Amazon and Rackspace
Indeed it's a nice problem to have :-) we are not yet there, we double every 3 months so far, so we still have 9 months to go :-)

however, getting out of cloud does not help. The problem is the high-availability. To maintain 99.9% we cannot rely on a single data-center, no matter what the claims on availability zones say. We have even seen network partition which are the worst of the worst. Once you go on the "Internet" there is no guarantees (unless dedicated lines).

solso··on Having fun with Redis Replication between Amazon and Rackspace
that's a good point... as mentioned in the post we use ssh tunnels and have experienced quite a few glitches, but in this case it has been running without a hiccup for 3 weeks (proof of nothing of course).

I suspect (guess) that not having an idle connection help the tunnel to not randomly drop.

solso··on Having fun with Redis Replication between Amazon and Rackspace
woops! @antirez in person :-) amazing job you have done with Redis.

Initially we didn't add compression to save bandwidth but to solve the network bandwidth issues between amazon and rackspace. However, it turns out that you can save on the bill quite a lot, in our case about ~ $700.

Regarding the binary, would be very nice to have. The first though is that compression would work better for us than binary since our keys are quite long (50/100 bytes) and very regular on their patterns, so compression works very well in our case. Probably can be extrapolated for other cases where the keys are typically heavier than the values.

← PreviousPage 2 of 2