Elasticsearch 2.0.0-rc1 released
elastic.co
elastic.co
2.0 is mostly cleanup, and doesn't really bring any big new features. The change that brings most utility is the merging of "filtering" and "querying" into a single query DSL. This maps better to how developers think about search, and reduces the mental overhead of having to decide if, say, a boolean operation should use a filter or query.
As I understand it, pipeline aggregations aren't strictly necessary -- you could do it client-side, so it's more of a convenience and optimization. Doc values are another optimization. This release is full of optimizations, cleanups and various low-visibility stuff of this kind. Few actual new features.
it's obviously not real pagination but the github issue explains why it isn't practical. If you know of other systems that have paginations of aggregations, perhaps you can reference those sources in the github issues and they can learn some tricks how to do it.
As for top hits, the "from" option doesn't let you do what you want?
My use case is that I have around +20M documents with non-unique hashes and each query should return an arbitrary amount of documents matching the query / filter as well as calculated meta-data based on the results in the aggregation field.
Now the issues is that if you want to have only one document per hash, you need to use a TermAggregation on the hash field followed by a TopHitsAggregation of size 1 to obtain the actual document rather than the hash field.
At this point you have many many buckets containing a single document but:
- you can't paginate them since TermAggregation doesn't let you do pagination for the reasons you explained and linked to above
- you can't calculate an aggregation on all the returned documents since all of them are in their own separate bucket (by hash)
It's also possible you could restructure your data to make it easier to extract the way you want.. I'm not sure if that's possible in your case, but many times in the NoSQL world, modeling your data in the best way for how you want to extract it is key for success.
This would mean that every time a user triggers this API call we have to insert those tens of thousands of documents into Redis and run an operation at the end. Even without taking into account the JSON serialisation / deserialisation costs, this would take forever where we need something that runs in around one or two seconds at the maximum.
The problem is actually pretty simple and can be summed up by this sentence: "I want to group documents by field X, take one document per group of X and return documents from index OFFSET to index OFFSET + LIMIT".
In MongoDB for example this would be quite easy with a $group + $first operation in the aggregation framework. Sadly MongoDB lacks the nice full-text search features that ElasticSearch has. It's highly possible though that going forward we'll have to hack a full-text-like search on MongoDB and switch from ElasticSearch to MongoDB since stuff like this doesn't seem to possible.
Postgres' text indexing isn't as advanced as ElasticSearch's (and indexing performance is definitely lower), but it's not bad at all. If you only need basic term support and not actual term vectors, you can use plain Postgres arrays, which have all sorts of supported operators, and can be translated into tables and back to perform some really neat in-memory queries.
One downside, of course, is that you don't get any sharding for free, and if your dataset is large enough you might have to manually partition your tables, either using Postgres partitions or by explicitly sharding (e.g., using pg_shard).
So why don't you use mongo your main data source and for the aggregations and use ES for the search (you just push changes to mongo when they update.) It's hard to beat ES/Solr/Lucene for text searches, but depending on what you need, Postgres' full text search may be fine. Most use-cases where search is the need use some other datastore and just ES for the text searches. Unless a text search is somehow part of the aggregation query..? in any event - good luck.
They were even kind enough to make an upgrade plugin: https://github.com/elastic/elasticsearch-migration
I find no mention about it in the release notes, so I'm guessing it's not... :\
Read through the series of pages to understand all major changes.
https://www.elastic.co/blog/elasticsearch-2.0.0.beta1-coming...