Auto Recovery for Amazon EC2
aws.amazon.com
aws.amazon.com
If you want something similar for ephemeral instances, do what we do: min 1 max 1 auto scaling groups. We've found that Amazon is pretty good at catching bad instances and terminating them, although on occasion we do have to terminate an instance manually. The autoscaling group takes care of the rest.
There is also a video of a talk about it: https://skillsmatter.com/skillscasts/6057-herding-cattle-wit...
We based our approach on Netflix but ended up building our own tools which we've now open sourced.
If you're only scaling up a few instances at a time then your number one consideration is probably making life easy for yourself.
edit: grammar
Examples of problems that cause system status checks to
fail include:
* Loss of network connectivity
* Loss of system power
* Software issues on the physical host
* Hardware issues on the physical host
All of these are on the physical host, which end users cannot control. So if AWS has an issue that kills your VM, if you don't have this setup then your instance is effectively dead?And there's no indication that the hardware and software "issues" are permanent or even fatal.
In some ways, I guess this answers my own question. Amazon doesn't know how long you might want to wait or if you have a VM that you would even want to have recovered, so configuring this lets you tell Amazon what your parameters around recovery should be.
Perhaps it's a co-promotion for CloudWatch. I would guess quite a few of their users had never heard of or seen a use for CloudWatch. Some of them might now enable "detailed monitoring" for $3.50/mo per instance while they're at it.
I recall several posts where a customer asked a question along the lines of "Why did my instance reboot" and then someone from Amazon replied something to the effect of "Sorry there was an issue with the underlying hardware but we did restart your server on new hardware".
This lands on the wrong side of pets-versus-cattle. AWS has been moving towards giving people what they want, but it's still best practice to use ephemeral storage and architect accordingly.
> If you've drunk the pets-vs-cattle koolaid
Generally drinking the kool-aid is a bad thing - any reason to favour pets?Its not worth the engineer time. Use EBS volumes, clean them up when they're no longer in use after termination. The only time you need local/ephemeral storage is swap or scratch space, or throughput you can't get from general or provisioned EBS.
Plus, you get auto recovery now without having to have architected for it ;)
The General Purpose SSDs and the Provisioned IOPS have more than handled our performance concerns. Since that last awful multi-AZ outage a year or two ago, there hasn't been much to deal with for us.
But sure, "Best PracticesTM".
This feature is currently available for the C3, C4, M3, R3, and T2 instance types running in the US East (Northern Virginia) region; we plan to make it available in other regions as quickly as possible. The instances must be running within a VPC, must use EBS-backed storage, but cannot be Dedicated Instances.
New AWS accounts can't even use EC2 Classic, it's effectively deprecated at this point.
Having this level of control can be nice, but most of it really needs to be optional because for most deployments it does nothing other than add an excessive amount of unneeded complexity.
Some of the APIs are outright hostile, e.g. 'delete_vpc' which makes you track down half a dozen dependencies (without providing hints about which those might be) before you're allowed to delete a VPC.
I've never gotten the impression that AWS is interested in building a 'cloud' for those less technically inclined. Heroku and others fill that void, EC2 is where you move after Heroku doesn't fit the bill and before dedicated hardware does.
[1] https://aws.amazon.com/blogs/aws/classiclink-private-communi...