Common Search – nonprofit search engine for the Web
about.commonsearch.org
about.commonsearch.org
Seems like right now they are focusing on getting contributors though.
It would be even better, if the web would be treated, at least for parts, as a digital library and non-profit organizations would recognize the value of access to such a resource and provide it (just as maintenance for roads or public schools).
We probably won't expose the underlying crawl data ourselves but being able to reference it just like [1] does is indeed important.
I can think of many other use cases as well, where a product needs to built from larger, but very selected set of raw input or pages.
BTW: Thanks for Hackday Paris 2011! Loved that event and venue :)
One problem with what you'd like to do is the pagination. Because we have to send queries to all the shards and then re-rank them, it becomes increasingly hard (and useless for most users) to build pages with p > ~20 (which is why all search engines heavily limit their pagination). So our main infrastructure definitely won't be optimized for that :/
However sending the top ~500 pages with a keyword + a domain filter should be doable pretty easily!
You do realize that you are talking about potentially a LOT of data?
To give you an example: The word "work" occurs on about 4% of all web-pages. So even if there were only about 2bn pages in an index, that would mean 80 million matching pages. Even if you only need their URLs that would be about 2.4gb of data assuming an average URL length of 30 bytes. Ok, compression can make that smaller, but still...
It would also mean that the server would need to make 80 million random reads to get the URLs. Even with SSDs that would take some time. Hmm, actually in this case it may be faster to just read all URL-data sequentially, than doing random reads. But in both cases we would be talking about minutes needed to get all that data from disk.
I currently have a search-index with about 1.2bn pages - I expect to reach 2bn pages by mid-May - that could be used to get the kind of data you need. But not in a realtime API. Not that amount of result-data.
> that could be used to get the kind of data you need.
Cool. Would you be interested in sharing or exchanging data?
What would be more useful to you, the raw data - meaning for each page a list of the keywords on it - or the reverse-word-index?
Raw-data may be better for batch-processing or running multiple queries at the same time.
My crawler currently outputs about 40-45gb of raw-data per day (about 30 million pages). Full crawl will be 2bn pages, updated every 2-3 months.
The reverse-word-index would be about 18gb per day for the same number of pages.
Reverse-word-index is already compressed, raw-data isn't.
There is a small problem with the crawl though, as it does not always handle non-ascii characters on pages correctly. I'm working on that.
BTW: I also currently have a list of about 8.5bn URLs from the crawl. About 600gb uncompressed. These are the links on the crawled pages. Obviously not all of those will end up being crawled.
Worst case we may build our own specialized crawler just for this purpose, but it would be nice if there was a useful search engine API we could leverage. And, of course, we'd be happy to pay for access to such an API.
Anybody here has any thoughts/experience with it?
If commonsearch can beat Google in that regard, then count me in. But I doubt it will.
I'm not sure if I buy that, but I do believe if we don't commit to alternatives it will be next to impossible for alternative search engines to get as good as Google with result quality and relevance. Google simply know too much about me and has performed so many more searches for a rival to outperform them. I still use DDG's 'g!' often but I feel like I'm doing my part to help DDG get better for me and other users.
Duckduckgo is now doing localized results (you can choose your region, so it's transparent, unlike Google's). The thing about Google is that, even if you're not signed in, it still tried to present you with personalized results (based on previous searches for that session, your IP, your region ... if you're searching from work; it probably factors that in as well).
When people talk about getting to the top of Google results, my response has always been, "Well you need to be more popular and relevant. Also you may be at the top..for some people, but not everyone."
I don't think search result quality is on a linear scale so it's hard to define "better".
The results will definitely be less personalized, which will be a big plus for some people, and a blocker for others. There will be a few other dimensions where we can stand out, and some where we will have a hard time catching up (index size for instance).
In the end, given enough contributors, I'm pretty sure the results can get "good enough" for most people, and hopefully "better" for some ;)
I suggest not to use AWS if you know that you'll need a server 24/7. Old-school hosters which offer dedicated servers are much cheaper for that use-case.
There are several offers here in Europe where you can get an i7-6700, 64gb RAM and 1tb SSD for less than €60/month. AWS would cost you at least 3-4x as much. You'll lose the flexibility of AWS, but save a ton of cash.
Isn't there more to the analysis than just comparing cpu before we can conclude it will save a lot of money?
It looks like their servers[1] use ~150TB source data that's already hosted on AWS disks. The source x.gz archives of the Common Crawl on AWS S3 are then imported to a Elasticsearch disks that are hosted on AWS.
To pull ~150TB of data using network speeds of 30 megabytes/sec[2] would take 60 days to transfer from AWS to another USA datacenter like Rackspace.
(Copying data from AWS to AWS isn't instantaneous either but it won't take ~60 days. At 60 days, the next crawl archive would have been released before you finished importing the previous one!)
Questions would be:
1) What are current 2016 network speeds between cloud providers?
2) What's the cost of ~150TB of network bandwidth?
3) From those datapoints, can we derive a rough rule-of-thumb where a certain amount of data exceeds the current capabilities (speed or economics) of the internet backbone available to projects like Common Search?
[1]https://about.commonsearch.org/developer/operations
[2]http://www.networkworld.com/article/2187021/cloud-computing/...
I'm pretty sure if you need to ingest ~150 TB you can pull it from AWS/S3 much faster than you think. To absorb ~150TB you'd need ~75 nodes. Given you can download partials of Common Crawl, you can break it up to 75 nodes downloading in parallel with 1gbit/s ports you should be able to pull it down relatively quickly compared to your estimation.
I'd bet you could pull ~150 megabytes/s [16 mbit/s per node].
http://commoncrawl.org/the-data/get-started/
> The Common Crawl dataset lives on Amazon S3 as part of the Amazon Public Datasets program. From Public Data Sets, you can download the files entirely free using HTTP or S3.
https://www.hetzner.de/en/hosting/produkte_rootserver/ex41s
> 2) What's the cost of ~150TB of network bandwidth?
Free.
> There are no charges for overage. We will permanently restrict the connection speed if more than 30 TB/month are used (the basis for calculation is for outgoing traffic only. Incoming and internal traffic is not calculated). Optionally, the limit can be permanently cancelled by committing to pay € 1.39 per additional TB used. Please see here for information on how to proceed.
> 3) From those datapoints, can we derive a rough rule-of-thumb where a certain amount of data exceeds the current capabilities (speed or economics) of the internet backbone available to projects like Common Search?
I suspect you are greatly overestimating the difficulties since most DCs basically let you ingest/download for free because of the asymmetry on their networks.
Regardless, I'll be keeping an eye on CommonSearch.
Sourcecode for deusu.org is on: https://github.com/MichaelSchoebel/DeuSu
OpenWebIndex is only in the idea stage as far as I know. Deusu has been running for over a year and already has about 1.2 billion pages in their index.
I think the next big thing in search will start as a niche thing. If you reduce your user domain you can have a better shot at producing better results even without tracking.
For example, you could develop a search engine for developers/IT people, make it's results better than google's and then expand to other domains.
I (like many other people on HN) use Google constantly when programming, and it's impossible to overstate the convenience and power of Google's almost creepy ability to guess exactly what the language and context of my search query is. It cuts precious seconds off of each query (when I am routinely make hundreds of queries per day), and more importantly cuts out the interruption of mental flow as you try to re-word your query into a format that the search engine will understand. This can potentially add up to hours of saved time per day, depending on how you calculate the impact of these features.
In order to duplicate this I don't think we can get around the need for "search profiles", which takes into account your location, interests, past searches, personal connections, etc, but it needs to be explicit and it needs to be my data. If I want to delete it or sell it, it really needs to be up to me. If we could figure out a secure way to do this, then we would have the framework necessary to compete with Google with an open platform. Until then, it's just not going to be worth it to switch for the vast majority of people.
An interesting metaphor for search engines' power is in
How many open source projects log more engineering hours than Google's search team? It's the flagship product of one of the largest corporations in the world.
For example, the project's data sources[1] says that the bulk of data comes from The Common Crawl. It looks like the CC is ~150 TB of data[2]. I'm not familiar with google.com internals but various sources estimate that their proprietary crawl dataset is more than a petabyte. (A googler could chime in here with more accurate data.)
So it's not as simple as the algorithm for Common Search being "more fair" than the algorithm for Google Inc. The underlying dataset in terms of quantity, recency, rules for the robot, etc all affect the algorithm.
This is not a criticism of the project. It is my attempt to understand what is not obvious on the surface level.
[1]https://about.commonsearch.org/data-sources
[2]http://commoncrawl.org/2015/12/november-2015-crawl-archive-n...
(I'm can't tell if each archive of MM/YYYY is cumulative or an addendum.)
I left Jamendo 6 years ago so I have limited influence on what they do now, unfortunately.
Common Search is a nonprofit and 100% open source so it is fundamentally different.
There is always a risk in trying something new. It might never pan out, or disappear in a puff of smoke.
Nothing would happen if nobody took a little risk.
If you're talking about privacy and transparency, it's better to operate in a place bound the European Charter of Fundamental Rights, rather than the US Constitution, because the former gives people much more rights with their data, how it's used, etc.
At scale, we'd probably have multiple legal entities in different countries anyway, like Wikimedia.
The explainer tool gives a really cool insight into the results: https://explain.commonsearch.org/
Sure it might not handle the long long tail but the top ten million searches would still be pretty useful.
We used http://www.discourse.org/, which I recommend.