AWS Elasticsearch Service Woes
havingatinker.uk
havingatinker.uk
Elastic (the company behind Elasticsearch) provides an AWS based SaaS Elasticsearch product https://www.elastic.co/cloud/as-a-service with a very sane version policy[2] that keeps pace with the latest Elasticsearch releases.
Disclaimer, I work for Elastic
1. https://aws.amazon.com/elasticsearch-service/faqs/
2. https://www.elastic.co/guide/en/cloud/current/version-policy...
I think it may be cheaper now though.
smart people cost a lot of money, having them on call at any time of the day or night costs even more money.
one of the interview questions i ask is "a customer calls in and said their system is unresponsive. what do you do?"
this will weed out a bullshitter 99% of the time, because the potential problem domain is the entire OSI stack, across the open internet, and also could be potentially on the customer side, which is most likely also a complex setup.
if you want someone who can actually fix the problem to help you, it's going to cost a lot of money. one of these people is six figures. now imagine a whole team of them, 3 shifts a day, every single day, forever.
Also I gave a poor example. I run a Es cluster so my thinking is at the point of my cluster is unavailable i know it's all red and it's not between my service and the endpoint.
if you feel that strongly about it, buy the support.
1. We were provisioned hardware that was less performant than promised (HDD rather than SSD-backed.)
2. The dashboard suggests significantly lower memory utilization than reality. We spent days going down the wrong rabbit holes, when we were getting timeouts simply related to out of memory errors.
3. When there's a shard failure, your only option is to reach out to support. You can't upgrade the database or anything.
We switched to Bonsai and it's been much more stable.
Secondly access control sucks. You can only whitelist by external IP, and you cant access elasticsearch from your VPCs directly, as they sit outside your VPC.
I abandoned it, wrote some chef and terraform and had a much more stable and flexible setup. However at an increased management overhead. If it works for you thats cool, but there are caveats, so beware.
We did get it to work in a useable manner by having the ES policy apply to a role (i.e. the principal is a role). If you than apply that role to your instances it will work for with instance profile based auth.
We are considering switching to the normal logstash-output-es plugin together with a AWS v4 auth signing proxy to make the setup more portable/less tied to AWS. I have a basic signing proxy setup working based on nginx/openresty. If you're interested just let me know.
Code: https://gist.github.com/nakedible-p/ad95dfb1c16e75af1ad5
Looks like it's been turned into an NPM-installable module too: https://github.com/santthosh/aws-es-kibana
Being stuck on an old version is painful but if your needs are simple then it's very useful. Ultimately if your needs are not simple then you probably should be running it yourself, just like anything else. Managed hosting for any database is never going to be as customizeable as you need if you operate at a huge scale.
The lack of transparency is my least favorite thing about AWS. I see the value in keeping brand new services or huge feature additions secret until they're released. But for small issues like when they'll update elasticsearch version or when they'll fix an acknowledged bug, I don't get why they can't share that information with customers (e.g. the way many open source projects track next version milestones on github).
Is it any different at the Enterprise support level?
What do you expect from AWS? They are part of a publicly traded company, they cannot make announcements willy nilly.
I personally feel once you've managed your own ES cluster for awhile, you tell yourself there's no freaking way any cloud system can abstract all these ugly (but important) configuration settings and corner cases into a neat, nice cloud product. No freaking way. I'm gonna manage this myself.
The recent versions are far more stable and includes lots of efficiency in terms of memory usage.
Elastic says they don't backport fixes either, so in this case, it might be a good idea to stay near the leading edge of the release cycle.
A roll-your own deployment is basically "install Elasticsearch, install cloud-aws, configure cloud-aws with a security group ID to identify other machines in the cluster, add that security group to all machines you want to cluster together, start the service". Totally braindead-easy.
You seem to have installation confused with administration.
Off the top of my head you forgot security, monitoring, logging config, backups, handling common production issues such as splitbrains, write multiplication, garbage collection snafus, upgrades between versions with questionably compatible internal apis.
Running any database in the cloud is non-trivial, and, running elasticsearch effectively in the cloud is harder than running Zookeeper or Cassandra, for example.
Source: I founded and sold an ElasticSearch hosting company.
My point is that the value-add from AWS' managed ES product is much, much smaller than it might be for other similar products because of the relative ease of administrating ES.
easy + clustered should never be in the same sentence.
Granted, none of us were proper ES admins, but we had a lot of experience working with system administration and specifically database performance and clustering. Despite that we were definitely in over our heads with ES.
Some overprovisioning will be required, but with the extra infra spend you're delaying the need for a dedicated role to manage it.
ElasticSearch is by far one of the most obnoxious software programs I ever have the misfortune to administer. I now avoid self hosting it at all costs.
It doesn't help that their logging (at least pre v2) is incredibly dense.
ES was pretty brittle in the pre-1.x days, but from 1.0 onward it's quite easy to work with. The logging is dense, but that's because it's thorough - a feature I really quite appreciate.
Can you explain this one in more detail from a technical point of view?
I've been managing out elasticsearch cluster for the past year and a bit. It grew from a single node that also ran kibana, logstash, and nginx, and stored our mysql backups to 9 data nodes, 1 client node, and a dedicated master. I have faced issues, but never had to rebuild the cluster. For the most part, reading up on the ES docs, and making a config change fixed the issue. Sometimes I've had to restart a node, but thats rare.
Amazon's typical approach for AWS products is to MVP them, and iterate functionality.
This often means that early services work very well for a handful of use cases. If your case deviates slightly in requirements, you will be better off in the short term running it yourself--assuming you can't change requirements.
In my experience, the AWS teams are very responsive to feedback and work pretty hard to add functionality to their mvp services.