Features of Solr vs. ElasticSearch
solr-vs-elasticsearch.com
solr-vs-elasticsearch.com
It's something I threw together in a couple hours, and figured I'd iterate and improve over the next couple days, so please bear with the mistakes.
I fixed the more glaring errors (copy field, dynamic fields , Django etc), and will continue to do so as comments come in.
Tire's docs are a bit lacking but it maps more-or-less 1:1 with ElasticSearch, so it's not too bad.
Very pleased with the performance ElasticSearch provides. The installation was a bit foreign for me personally (a java service container? openjdk 6 or 7?) but it's lightning quick and very flexible.
I chose to go the ElasticSearch route mainly due to my need for Geospatial indexing. Solr does it too, but Foursquare[1] uses ElasticSearch and so that got me interested in learning more. Geo queries are really fast. My last experience with Geo involved GeoDjango and all kinds of obnoxious hacks to PostgreSQL to make it work. With ElasticSearch you tell it to index a point and boom you're off to the races.
[0]: http://freshbsd.org/
It'd be great if their docs had versioning (like the Apache HTTP Server Project), but I suspect that isn't on anyone's roadmap.
That said, they carry a lot of annoying edge cases when determining the adjacency of quadrants, so they're hardly a panacea to geospatial search. Lucene 3 and 4 make a lot of progress in spatial search, but there's still a fair bit of room for improvement.
That's because R-Trees don't scale well with random write loads.
R-tree insertion performance is extremely dependent on insertion order (search for "sort tile recursive"). They're best used for problems where the data can be bulk-loaded and left alone. If random writes are an important part of the problem (as they are for most web-based tools), R-Trees are a bad idea.
I can't speak for ElasticSearch but there are a couple things in the Solr list that I'm not sure about.
"Multiple document types per schema" - You can use dynamic fields so that you don't even need to define your document schema
"Schema change requires restart" - I think in MultiCore it happens when you swap cores (which is a good way of running solr) [0]
[0] http://stackoverflow.com/questions/10417422/solr-schema-chan...
Personally, I'd like to see similar comparisons for other search engines, like Sphinx and Postgres Full-Text. When I talk to people about search engines, the first questions they ask me are to compare one against some other.
I almost wish it was more of a standard data store. Here's to hoping RethinkDB can fill that void.
I setup a cluster of about 15 cc2.8xlarge machines (5 Shards with 3 replicas each) containing 240Gb worth of documents (48gb per shard). Each node was given on the order of 40GB heap space. While performing load tests with a relatively small load (~150 QPS) after a few minutes the garbage collector on nodes would kick in and run on the order of 15 to 30s. This had a cascading effect of causing zookeper to think nodes were down, start leader re-election, etc.
Admittedly I am quite inexperienced when it comes to dealing with applications using such large heap sizes. Though I tried a few different JVM options with respect to GC I was unsuccessful in resolving the problem.
If any folks here happen to have some good resources regarding GC and large Solr clusters I would definitely be interested.
Jesus dude. I couldn't reproduce that with my single or multi-node ElasticSearch clusters if I wanted to.
How were the EBS backing stores setup for these EC2 nodes?
Edit: Also, when I was talking about "them" falling apart, I meant Postgres or Sphinx, not Solr/ElasticSearch.
Well-configured Solr and ElasticSearch clusters can work very well for most people.
Try it again with sane GC parameters, e.g.:
-Xmx<N>G -Xms<N>g -XX:NewSize=<N/2>G -XX:MaxNewSize=<N/2>G -XX:+UseConcMarkSweepGC -XX:+DisableExplicitGC -verbose:gc -XX:+PrintGCTimeStamps -XX:+PrintGCDetails -XX:+CMSIncrementalMode
Where <N> is a value between 2-8.Edit: I was benchmarking a similarly sized (though very differently configured) Solr cluster for a well-known internet company, and was able to tune it to do 5000qps, with p50 ~2ms and p99 ~20ms.
I was starting to think that since the heap space was so big perhaps I should be worrying about page sizes as well. While I tried various GC settings (UseConcMarkSweepGC, ConcGCThreads, UseG1GC, etc. ) I didn't take a stab at playing with the size of New Genearation. Could you explain the reasoning behind this? Is the idea that most objects die young so try to increase the number of short run minor GCs and avoid bigger Major GCs? I am quite interested.
Edit: Regarding the cluster you were working on. Would you be able to give general dimensions to the number of nodes & partitions in your cluster + memory for each? Just trying to get a general guideline to aim for.
In general, you should have enough unallocated memory on the box to cover your working dataset (it'll get used by caches and memmaps). If you can, find a way to exploit data locality. I shoot for (number of cores * 1-4)-ish partitions per box depending on workload. Using bigger boxes is usually better, because you can avoid communication latency and variance that arises from having tons of boxes.
If you want to know more, you can email me at kyle@onemorecloud.com.
If, as was my case, the most complex search you've ever made before was a fulltext search on a database field then you'll be lost for a good couple days until you understand what's going on.
That means that as new documents arrive with fields that were never seen before, if those fields end in _ts, they will be properly indexed.
That feels like dynamic field support to me.
http://elasticsearch-users.115913.n3.nabble.com/Apply-dynami...
I'm skeptical about claims of different quality of relevancy results, since IndexTank/Searchify is also based on Lucene (looks like 3.0.1 in the canonical repo), and should share all the same fundamental relevancy and scoring functionality.
It's actually heavily tweaked (this is what the IndexTank team has told me), and apparently contains components of Solr