Major Changes from Solr 4 to Solr 5
cwiki.apache.org
cwiki.apache.org
fruit extractor => fruit juicer, citrus juicer
Does anyone experienced enough have a clue if SOLR 5 can help with that?[0] - http://opensourceconnections.com/blog/2013/10/27/why-is-mult...
https://github.com/healthonnet/hon-lucene-synonyms
The Solr guys don't give a flying F about this issue though
I've tried things like this
http://community.zimbra.com/documentation/w/documentation/se...
repeatedly. They never seem to work right.
Have Solr listen on localhost and have your web app talk to Solr. If your Solr is visible to the world, you're doing it wrong.
Edit: by saying that it's not a web-based application I mean that it shouldn't be on teh interwebz -- it's obviously a webapp in the sense that it mostly speaks HTTP.
I agree however that SOLR is best off doing one thing well, web page security can be implemented e.g. by Apache.
Or to whichever friendly consultant you decide to hire to help out with that. Wink wink. Nudge nudge.
I don't think there's anything out of the box in Solr 5.0 that changes that. SOLR-4470[0] should be able to do that, but it hasn't been committed. Apache Sentry[1] adds role based access control to Solr, but it's only been tested up to Solr 4.10 and with kerberos (not basic password protection). It comes nicely integrated out-of-the-box with Solr as part of Cloudera Search[2]; otherwise, you'll have to do some manual setup to get it to work.
[0] https://issues.apache.org/jira/browse/SOLR-4470 [1] https://sentry.incubator.apache.org/ [2] http://www.cloudera.com/content/cloudera/en/products-and-ser...
I also did a Meetup on this just this week http://www.slideshare.net/detnavillus/the-well-tempered-sear...
Check out slide 18 - autophrasing + synonyms: Precision 100% recall 100% Bag of words OOTB Solr/Lucene NOT so!
The code is on github and is a Lucene TokenFilter so it should work. I used 4.10.3 for the Meetup demo
Solr includes Admin UI console which is free for production, ES has one that's only free for development.
Solr has contributors from a lot more different companies, so grows into multiple directions at once.
If you want to compare on a technical level, you can see my presentation from the Lucene/Solr Revolution back in November: http://www.slideshare.net/arafalov/solr-vs-elasticsearch-cas...
Solr is completely free. If that's not an issue for you and you are ready to pay, then you should compare Elasticsearch to LucidWorks Fusion, not directly to Solr.
Scaling-wise? Distributed Elasticsearch doesn't have a Zookeeper dependency, which is nice. But Solr has more sharding flexibility and partition tolerance.
Like everything else, depends on your needs.
Yes, see [1] and [2] for a comparison between Solr and Elasticsearch. I should mention also that ES are working on the partition tolerance issue as described in [3]. I am currently using 1.4.3 and am wondering if anything was addressed in 1.4 for resiliency.
[1] http://lucidworks.com/blog/call-maybe-solrcloud-jepsen-flaky... [2] https://aphyr.com/posts/317-call-me-maybe-elasticsearch [3] http://www.elasticsearch.org/blog/resiliency-elasticsearch/
Speaking for myself I did that and found that SOLR was a lot more performant. I needed a high-traffic solution without a bunch of servers.
I find that ES tries to do too much with all the dashboards and monitoring etc.
SOLR keeps it simple and thereby does not incur the performance penalty.
Also when an ES cluster goes south (cluster health "yellow" or "red") it seemed like a pain to troubleshoot and determine the real reason WHY. SOLR seems more durable and when something needs investigation you get a clear message in the log.
If you are starting something new and just don't know your traffic requirements and are scared of it "going viral" like a hot new mobile game ES may work for you though. The one thing it excels at is adding more ES servers to the cluster quickly. So if you need more servers ASAP and don't care about the cost ES has that covered.
You save money and can do a reindex very quickly.
A basic "reindex" command would be a cool feature though.
Neither Solr nor Elasticsearch should be treated as a primary data store, so that's probably why reindex-in-place is not the highest priority.
There is no technical reason why Solr/ES could not do diff-based indexing. ES (I don't know about Solr) admittedly uses a single Lucene index per logical index, so changing a single field mapping involves reindexing the whole index, not just that one mapping.
But if the mappings were properly versioned ES could simply create a new version, index everything (from the original contents), and then swap. Locking the original index should be a non-issue.
It's a matter of who does the releasing (i.e. curating of patches, responsibility for building community, etc.): Solr is governed by the Apache Software Foundation and associated volunteers, ElasticSearch is governed by the private entity.
IMO (and this is an increasingly rare opinion) the license is the most important piece and they're both under the ASL 2.0.