Elassandra: Elasticsearch implemented on top of Cassandra
github.com
github.com
It was really interesting work.
OP - was elassandra influenced by either of those in anyway?
[1] https://github.com/tjake/Lucandra [2] https://github.com/tjake/Solandra
[1] http://nicklothian.com/blog/2009/10/27/solr-cassandra-soland...
Therefore Stratio's Cassandra Lucene Index is worth an equal mention. https://github.com/Stratio/cassandra-lucene-index
Elasticsearch is terribly broken in a lot of ways, but it's awfully easy to get up and running.
I'm not sure which I'm holding out for. That Solr will be easier to manage or that ElasticSearch will stop doing terrible things.
When you spend lots of time fine-tuning the exact combination of filters, ranking algorithms, weighting of fields and so on, the api just completely switching certain methods kind of hurts.
Trying out Elasticsearch, my experience was that it really wants to be run in a cluster, but it also loses data pretty easily. I had more issues with it crashing and it's generally a lot hungrier for memory.
Both have non-obvious shortcomings. Solr's schema will make you believe that it likes deeply nested JSON documents. False! It actually wants pretty flat "documents" without nesting (you /can/ nest, but it usually doesn't do what you want without some extra legwork). ES will have you believe that it supports lots of query types and they'll all perform great on semi-structured data. My experience was that it was difficult to predict performance, but that generally the fancier the query the worse it would be.
Solr's querying functionality is not extremely powerful (though they "helpfully" made it offensively complex with different query parsers and stuff) but performance has always been excellent for me.
IMO, if you don't need clustering, Solr is definitely better. A cranky but robust piece of engineering from before scaling was everything. ES has better documentation, a better "getting started" story, and is generally a lot more user-friendly. Aphyr's posts about it have made me wary of using it without a re-indexing story.
I haven't tried Solr's scaling stuff because I haven't needed it, but I would expect it to be in pretty rough shape compared to ES because it's not a primary use case for Solr and it is for ES.
The new SQL / Parallel Streaming has also made querying multiple collections a cinch.
Most of the problems that I have had are with the ZK clients and not the server. As long as you follow the operational documentation (there are a few basic rules) it hums along nicely. We have a few clusters with Kafka and have a decent process in Chef:
On the other hand, it's had its share of improperly handled split brain scenarios. I still think it has problems with partial split brain (where A and B can't talk, but C can talk to both of them).
Zookeeper as a hard requirement feels like asking me to supply an entire container port to unload one container from a semi-trailer.
Thanks for sharing the cookbook :) . Do you use Zookeeper for anything other than Kafka?
I just clustered ES with Docker somewhat recently. It was pretty good at discovery I must say. Well, I'm no great search engineer I just throw these things up when I need 'em so good to know all this. Thanks again.
> This is probably not true: Crate 0.54.9 uses Elasticsearch 1.7, which definitely loses updates. If Crate loses updates, it’s unlikely that you can guarantee reading the latest write, let alone reading that write ever again. Even the unreleased Elasticsearch 5.0.0 still fails its Jepsen tests according to Elastic, so claiming linearizable reads on single keys might be a bit of a stretch.
https://aphyr.com/posts/332-jepsen-crate-0-54-9-version-dive...
1. It mentions using secondary indexes - its my understanding thats a huge no-no, as they have to hit the whole cluster 2. Uses "lightweight" transactions - also another perf hit, as lightweight transactions have (anecdotally) a 6x slowdown...
I like the idea but I'm curious if these are issues and whether these uses are something the author is looking to replace...
Very interesting idea though!
(CoAuthor of cassieq here so these were things we had to learn about.)
lightweight transactions are only used on schema-changes (which are/should-be rare)
As for the indexes - are they standard Cassandra secondary indexes? "Custom secondary indexes" - does that mean that it just looks like a secondary index, but is actually backed by Elastic search?
Though you can't query it from cassandra yet. You have to use the elastic-search rest-api.
Would like to know more about how indexing is handled.
* Cassandra update are automatically indexed in Elasticsearch.
* Full-Text and spatial search on your cassandra data.
* Real-time aggregation (does not require Spark or Hadoop to group by)
* Provide search on multiple keyspace and tables in one query.
* Provide automatic schema creation and support nested document using
User Defined Types.
* Provide a read/write JSON REST access to cassandra data (for indexed data) - SASI requires 2 passes on disk to fetch data: 1 pass to read the index files and 1 pass for the normal
Cassandra read path whereas search engines retrieves the result in a single pass (DSE Search has a singlePass option too).
By laws of physics, SASI will always be slower, even if we improve the sequential read path in Cassandra
- Although SASI allows full text search with tokenization and CONTAINS mode, there is no scoring applied
to matched terms SASI returns result in token range order, which can be considered as random order from the
user point of view. It is not possible to ask for total ordering of the result, even when LIMIT clause is used.
Search engines don't have this limitation
- last but not least, it is not possible to perform aggregation (or faceting) with SASI.
The GROUP BY clause may be introduced into CQL in a near future but it is done on Cassandra side,
there is no pre-aggregation possible on SASI terms that can help speeding up aggregation queries