Deploying Elasticsearch on a 150 node cluster to index 10B documents
spinn3r.com
spinn3r.com
ES 2.3 brings the reindex API [0], which is an absolute godsend. Also, that is a remarkable amount of computronium brought to bear on what I would charitably describe as a moderately-sized dataset. Is the 500ms query time an absolute hard requirement?
And I may be admitting my ignorance here, but there's also this statement:
"However, Elasticsearch doesn't have a way to efficiently tell us how many documents were returned in a given response."
Are you looking for something above & beyond the "hits" value returned in a query response? Or am I missing something? ex:
{
"responses": [
{
"took": 16,
"timed_out": false,
"_shards": {
"total": 3,
"successful": 3,
"failed": 0
},
"hits": {
"total": 38,
"max_score": null,
"hits": [
{ ...
[0]: https://www.elastic.co/guide/en/elasticsearch/reference/2.3/...If you do want to return all results, you can use the scroll API.
There's no guessing with the results in Elasticsearch if you read the (surprisingly accurate) documentation.
Edit (moved from reply to myself): I guess I should clarify before someone corrects me: you could get fewer results than the size parameter. But in that case the total hits should be less than your size parameter and it should be clear that the number of returned documents is the difference.
Although, to be fair, the related issue[2] specifically asks for an HTTP header.
JSON parsing is generally pretty fast, but I can understand it being unnecessary overhead (although in general, there are other more significant opportunities for optimization).
If you've ever implemented a JSON stream reader/writer, then you've done partial JSON parsing. Check out the tool `jq`[2], which will parse partial JSON surprisingly well and fast (and this is just a utility).
[1] https://github.com/elastic/elasticsearch/issues/18312
[1] https://www.elastic.co/guide/en/elasticsearch/reference/curr...
I was recently at an Elasticsearch meetup hosted by HomeAway and pleased to see that may of the patterns that emerged at Bazaarvoice were generally repeated in HomeAway's deployment. I definitely encourage anyone considering Elasticsearch to explore its utility for more than just logging. So many times I read about or hear folks using Elasticsearch to index "billions of documents!!" only to find that only a few million documents are actually in an open index.
Does anyone know of any heavy users of Elasticsearch for non-logging workloads in AWS? So far it seems like most of these folks are running out of a colo or their own data center. I've been running Elasticsearch in AWS for a few years now and simply can't imagine dealing with the inconvenience of managing the hardware down to the specific number of nodes. (If you asked me, I'd say I don't know how many nodes -- lots!)
Why do they have to bill synchronously (in realtime)? Why not return the response, then calculate the fee asynchronously? The customer is probably being billed at discrete intervals (e.g., weekly or monthly) anyway.
Much easier if it's just an integer header field.
Disclosure: I am VP, Product Strategy at MarkLogic...