The Architecture of a Large-Scale Web Search Engine, Circa 2019
0x65.dev
0x65.dev
[1] https://0x65.dev/blog/2019-12-01/the-world-needs-cliqz-the-w...
How we collect data : https://www.0x65.dev/blog/2019-12-03/human-web-collecting-da...
How we build the search using this data: https://www.0x65.dev/blog/2019-12-06/building-a-search-engin...
Feel free to peruse these posts and ask questions!
Are you planning to create strong tools in that area ?
For example: custom search engines, the NEAR operator, limit search to sites that don't update that often or aren't linked to very strong sites(against SEO), etc
https://old.reddit.com/r/firefox/comments/74yo19/cliqz_and_m...
There were more recent discussions about Cliqz no latter than this month, in particular here: https://news.ycombinator.com/item?id=21676252
[disclaimer: I work at Cliqz]
> We have been posting multiple articles on our tech blog, explaining what we do and how we do it in great details.
it's possible to have both great tech and loose morals - the two are not mutually exclusive, and one does not absolve the other (e.g. facebook's social experiments)
has there been a followup to any of the points brought up in the reddit thread?
There is plenty of documentation on data collected (see first posts regarding Human Web on the tech blog), how anonymization works, why record-linkability on data collected is prevented (and forbidden), etc. Furthermore, source code can be inspected, as well as traffic in the case documentation is not enough. I believe that is a better proxy to assess "morality" than random accusations on reddit or opinions formed solely on a half-baked press releases.
Do we need to refute all miss-conceptions and FUD that might arise due to the fact that 1) we collect data to build our services (search) and 2) we are funded by a media company (VCs seem to be more pure for an unknown reason).
The answer is no. Cannot recall who said that it takes much more effort to refute BS than to generate it. (That does not go for your comment in particular, that's why we replied, but for many of the comments and some of content of the subredit that you mention.)
Question is, is it opt-in data collection or do you make the choice for me? If it's opt-in, great. Otherwise, I don't want to read your "plenty of documentation" and so on and so forth.
[1] https://vespa.ai
Work on Cliqz Search started way earlier ~2013. Our work on Kubernetes and modernizing our architecture was also started around year 2016.
[1] https://www.verizonmedia.com/press/open-sourcing-vespa-yahoo...
This blog post, and other in the series, mention RocksDB is used for the index, but it's not explicitly described how and to what end. I 'd love to know the details.
Note that kubeflow R&D is all done by google cloud.