Elastic: Several customer facing deployments deleted across multiple regions
status.elastic.co
status.elastic.co
These safety layers help present destructive operations more visibly before they are completed.
For destructive actions, putting a few days between “take server offline” and “throw disk into the shredder” often is possible, too.
But yeah, how can you run mission critical if most cloud services have SLAs two or three nines weaker than needed?
My most critical infrastructure for my one-man SAAS is all third party infrastructure run by large companies. My non-critical infrastructure is self managed for cost savings.
>how can you run mission critical if most cloud services have SLAs two or three nines weaker than needed
That's exactly it. Cloud providers usually provide SLAs in the range of 95-99%. Amazon doesn't provide a full refund until monthly uptime goes below 95 %. Elastic apparently doesn't provide an uptime guarantee at all, they only provide an SLA for support ticket response times. And only on gold and platinum subscriptions.
This incident lasted for more than 24 hours («most» instances restored after 22 hours). It doesn't matter if it's someone else that has to wake up at 3 am to fix the issue, when they're unable to fix it within reasonable time. Mission critical apps simply can't be down for 24 hours. And it's fully possible to design HA elasticsearch deployments. Elastic Cloud just isn't one of them.