Lessons Learned from a Year of Elasticsearch in Production
tech.scrunch.com
tech.scrunch.com
Internal operational metrics like thread pool size are excellent for diagnosing failures to reach target latencies, but keeping thread pool queues low shouldn't be the direct goal.
We use the AWS ElasticSearch service, and while it does not provide the same flexibility, it's been fantastic for us. No configuration files, no weird restart behaviour, no thread pool madness - it just runs. You can upgrade your instance sizes seamlessly and it integrates well with other AWS parts (ex: cloudwatch).
There are a few hosted elasticsearch services available (Elastic Cloud, AWS ES, Qbox, etc) which would likely be a better option for the team that needed a hands-off elasticsearch cluster.
Regarding those examples:
1) For RHEL and Windows, paying is not optional 2) Mongo is so buggy most people need to pay for support in production 3) MySQL works well, but don't expect anything more than "forum answers" from Oracle support. Only 1 in 1000 companies paid for support according to the former president of MySQL AB.
Actually, companies are reluctant to pay for support for Open Source products, hence the business model tweaks of Redhat and the Open Core models of Nagios and other smaller Open Source vendors.
The tradeoff in value wasn't worth the time it was requiring, so we outsourced to compose.io -- they've been a good fit price & stability-wise, and feel pretty close to a blackbox for the benefit. If you need a large cluster however, you'll pay handsomely with any of the SaaS providers when last I looked at it. For large quantity work we've moved to AWS Redshift.
I should add that all "hot" data fits in 24g of memory.
I have two clusters running.
Largest one is an ELK stack. 7 nodes (soon to be 8, we're logging too much :) ), no dedicated master nodes, ~3TBs of data (1.5TB data x 2 for the back up). Need to make sure they are split appropriately into availability zones in case you lose a cloud zone. Need to have the right # of minimum master nodes otherwise you can hit a split brain issue. 500GB SSD, 4vCPUs, 26gigs memory. It eats ~5 gigs an hour peak load without breaking a sweat. Can basically take any heavy query people have thrown at it kibana wise.
The cluster is much smaller than others---having 12 nodes eating 32GB memory a piece, and handling only 50GB of data. I run a consumer-facing website from it, where all queries hit the ES cluster with limited caching at a pace of ~8,000 qps at peak for fairly heavy computation aggregations.
ES scaling has consistently been simple as adding another box ever 3-6 months.