Amazon AWS had a power failure, their backup generators failed
twitter.com
twitter.com
I haven’t seen a report yet on exactly why their generator failed, but from what I’ve heard, the power failed, and the backup generator kicked in and ran fine for over an hour, but then it failed. This sort of thing sucks to deal with but it’s also inevitable. Of their hundreds of datacenters, a mechanical failure is going to happen occasionally no matter how good their maintenance plans are.
So the key when using cloud services like AWS is to plan for the possibility of failure. EBS expects an annual failure rate of 0.1%. So one out of a thousand EBS volumes will fail in a given year. If you operate at the scale of thousands of servers in AWS, you see this sort of thing all the time. Luckily, EBS also makes it trivial to take volume snapshots which are stored in S3, which has much much higher reliability and durability. So if you have data in EBS that needs to be kept safe, take regular snapshots. Here’s a doc that explains how you can set up scheduled, auto-rotated snapshots: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/snapshot...
[01:30 PM PDT] At 4:33 AM PDT one of ten data centers in one of the six Availability Zones in the US-EAST-1 Region saw a failure of utility power. Our backup generators came online immediately but began failing at around 6:00 AM PDT. This impacted 7.5% of EC2 instances and EBS volumes in the Availability Zone. Power was fully restored to the impacted data center at 7:45 AM PDT. By 10:45 AM PDT, all but 1% of instances had been recovered, and by 12:30 PM PDT only 0.5% of instances remained impaired. Since the beginning of the impact, we have been working to recover the remaining instances and volumes. A small number of remaining instances and volumes are hosted on hardware which was adversely affected by the loss of power. We continue to work to recover all affected instances and volumes and will be communicating to the remaining impacted customers via the Personal Health Dashboard. For immediate recovery, we recommend replacing any remaining affected instances or volumes if possible.
HA would normally use two AZs, 21C architectures would use three AZs, and other patterns such as pilot light let you use additional AZs without a significant cost hit. Further, when you can make workload-handling instances so much smaller (even down into the T sizes, and with cattle patterns you can start to leverage spot pricing), each additional AZ you add to the 21C mix represents a smaller percentage of your capacity lost in an AZ outage.
What factors have you choosing an automation powered CSP such as AWS but using only a single AZ out of a half dozen?
If using a single AZ and tense during outages, why not use Hetzner, Softlayer (now IBM Cloud), etc.?
Either way, GP is using obscure terminology at best.
> The term pilot light is often used to describe a DR scenario in which a minimal version of an environment is always running in the cloud. The idea of the pilot light is an analogy that comes from the gas heater. In a gas heater, a small flame that’s always on can quickly ignite the entire furnace to heat up a house
Sure, I’m being nitpicky, but that’s a terrible explanation of the purpose of a pilot light in a legacy furnace.
I would like to note that the response time suggested by the tweet is a bit bad, 4 days to realise something and send out response / alerts to customers is a bit slow even for amazon. But then again, everything goes slower for bigger things, and amazon is quite big i'd say. not sure what the SLA response time to such an incident is, so it might be within agreed times...
When a datacenter loses power like this, a few of the storage arrays will just not come back online. But another few will take time to run through their corruption recovery process, and it may take a long time and some service by a human (eg parts replacement, etc) before they can be certain a particular volume is not recoverable. Given their scale and the timing, at the beginning of a holiday weekend, four days is annoying, but not bad.
Therefore, if you were impacted, its your fault, not AWS. Sorry.
https://12factor.net/ can help guide
[1]: https://stackoverflow.com/questions/13576363/does-taking-a-s...
[2]: Yes it's a big if but S3 for durability against data loss is like the US Treasury for risk free returns on T-Bills.
The EBS volume will still fail but you can restore it from the higher reliability S3-backed EBS snapshot.
Regardless, it's important to be aware of the risks of whatever tool you're using. It's unrealistic to expect any provider to be able to avoid failure entirely. You have to be aware of possible failure scenarios and have your own plan to address them.
It is inevitable that systems will fail. The best the industry can do is work to reduce the number of failures and understand failure modes well so that they can be planned for. AWS does a very good job of this in my experience.
Regardless of whether applications are hosted in the cloud, on premises, co-located or in some hybrid configuration, it's important to design for that inevitable failure and keep business decision makers in the loop while doing so. Understanding requirements around RPO and RTO are extremely important in developing an architecture which meets the needs of the business, yet is still cost effective.
Likewise, I guess it's easy to speculate from the peanut gallery, but I wouldn't be surprised if the backup generators just hadn't been sufficiently tested and maintained because, well, they're backup generators.
AWS says they plan to install a backup fuel delivery system at this particular datacenter (it wasn't detailed whether this SPOF is common at other datacenters or if this was an outlier) and that they have already updated their notifications to be more aggressive.
"Your nines are not my nines" - https://rachelbythebay.com/w/2019/07/15/giant/
For how many users was this 100% of their business?
Single-digit outage percentages for cloud services like AWS look like no big deal from the big perspective, which means that they often aren't a high priority. But when you're one of the customers in the 2-3%, it is a big deal, and you want it to be a high priority.
Small hosting providers may have only 3 9s of availability, but when your site goes down their world stops until it's fixed. I've seen reports of 5 9s of availability from Amazon, but when your site goes down, their alerts are still all green and they'll call you back at their leisure.
… but they still chose to ignore all of the prominent warnings and architectural guidance, not to mention avoiding use of the services which have HA built-in. I mean, I'm sympathetic to anyone who had a bad day with a forced learning experience but it's not like this is some dark secret.
So that's it, blame the user and caveat emptor?
> not to mention avoiding use of the services which have HA built-in.
Many (all?) of these services tightly couple you to Amazon, so avoiding them is a very reasonable decision.
If I sell you a loaf of bread and you complain that it's not a sandwich, is it anything else?
> Many (all?) of these services tightly couple you to Amazon, so avoiding them is a very reasonable decision.
That's just a cop-out: checking the “multi-AZ” box in RDS completely avoided this problem with zero lock-in. If you're deploying containers, you have multiple options which are portable and avoid this completely. If you're deploying EC2 instances, again you have options with very limited lock-in (e.g. auto-scaling with multiple AZs).
More importantly, that's also a business decision: if you're that worried about lock-in you are accepting responsibility to operate the alternatives. For example, following industry-standard practice might suggest that you run everything in Kubernetes in multiple AZs, regions, or providers but it would never support running everything in a single AZ.
It's managed infrastructure not some miraculous alternative universe where probabilities do not apply to you.
From the docs:
"Amazon EBS volumes are designed for an annual failure rate (AFR) of between 0.1% - 0.2% ..."
The morning of the worst of the storm, we completely lose access to all services at that facility. Super unusual, everything we have there is redundant. So I do some minor investigation, and get on the horn to them.
They were being super cagey. "Hey, we lost all access to our systems." "Ok, I'll open a ticket and we will investigate." "Uhhh. It feels like it's a big problem with the data center, are you guys having problems or is it just us?" "I can't say anything more until we've completed an investigation." "I'm trying to decide if we need to start failing over to our DR site, or if I need to put chains on the truck to drive 50 miles to the data center in this storm. Is anyone else having problems? Are fire alarms going off?" "We have received multiple reports of problems."
Power was back on in less than half an hour, but they still weren't saying anything for a few hours. Spent that time trying to figure out if we should shut everything down and ride it out, or if they were back in business. We had one system that suffered disk corruption, despite having a (according to the weekly testing) correctly operating BBU on the RAID.
So what happened? It shouldn't have been possible, our cabinet was being fed by two lines from independent PDUs. Each PDU is fed by 2 independent UPSes (one shared between the two PDUs), each UPS fed by a dedicated generator. Should have required 3 failures to bring us down.
They eventually reveal that they had had one of their UPSes down for weeks, waiting for replacement parts. The other two UPSes had independent failures (one was a controller board, one was battery related). They said they still did quarterly full load tests of the power systems, but reading between the lines I think they weren't testing these two UPSes because the other one was not there to back them up.
Still, one power event in 15 years isn't too shabby.
I know of a large company that had the data center emergency cutoff button next to the automatic doors on the way out. Sure enough, a contractor hit it one day thinking it was the way to open the doors.
It doesn’t help that ebs backup takes forever, especially initially.
I'd rather just put up with the agony of RDS to get a point in time restore and treat my instances' data as volatile.
It was a case of either lying or incompetence. Hunt didn't like that and blocked me for calling it out.
Link?
It’s annoying when amazon has outages but they have local outages all the time and they give you all the tools you need to handle them.
>>Then it took them four days to figure this out and tell us about it.
Is there some post mortem that just came out?
>>We want to give you more information on progress at this point, and what we know about the event. At 4:33 AM PDT one of 10 datacenters in one of the 6 Availability Zones in the US-EAST-1 Region saw a failure of utility power. Backup generators came online immediately, but for reasons we are still investigating, began quickly failing at around 6:00 AM PDT. This resulted in 7.5% of all instances in that Availability Zone failing by 6:10 AM PDT. Over the last few hours we have recovered most...
https://www.datacenterknowledge.com/archives/2017/04/07/how-...
> The piece of technology Amazon designed to avoid this type of outage is the firmware that decides what electrical switchgear should do when a data center loses utility power. Typical vendor firmware prioritizes preventing damage to expensive backup generators over preventing a full data center outage, according to Hamilton. Amazon (and probably most other large-scale data center operators) prefers risking the loss of a sub-$1 million piece of equipment rather than risking widespread application downtime.
> When everything happens as expected during a utility outage (which is the case most of the time), the switchgear waits a few seconds in case utility power comes back (also the most common scenario) and if it doesn’t, the switchgear fires up generators, while the data center runs on energy stored by UPS systems. Once the generators are stabilized, the switchgear makes them the primary source of power to the IT systems.
> Last year’s Delta data center outage was attributed to switchgear “locking out” the generators at the airline’s facility in Atlanta. That’s what most switchgear is designed to do when it senses a major voltage anomaly either in the data center or on the incoming utility feed. Plugging a live generator into a shorted circuit will usually fry the generator, and switchgear locks generators out to avoid that.
Sometimes instances just go out to lunch. Sometimes an AZ goes down. Chaos Monkey isn't just a good idea, it's required for reliable operation.
But please, please, if you are going to treat AWS like a VPS, at least don't do it in us-east-1! It seems to have more outages.
Our setup related to instances and EBS includes: At least 2 instances in different AZs, an ELB in front of them, a backup running at our hosting facility (though this could just be a different AWS zone, or different provider), and DNS with full-paper-path health checks that switch DNS over to the colo servers if any component of the primary fails.
Cloud vendors, of course, get to hide all that from you until something goes wrong.
It seems that it's pretty hard to get this right from the beginning too, every large datacenter ends up learning this again after they have a sequence of power incidents. That said, successful switch to generator for 1 hour and then generator failure is not a terrible outcome; if there was a notification, that's enough time to evacuate critical systems (assuming you have a plan).
The headline sounded so clickbaity that I ignored it before seeing this thread
[01:30 PM PDT] At 4:33 AM PDT one of ten data centers in one of the six Availability Zones in the US-EAST-1 Region saw a failure of utility power. Our backup generators came online immediately but began failing at around 6:00 AM PDT. This impacted 7.5% of EC2 instances and EBS volumes in the Availability Zone. Power was fully restored to the impacted data center at 7:45 AM PDT. By 10:45 AM PDT, all but 1% of instances had been recovered, and by 12:30 PM PDT only 0.5% of instances remained impaired. Since the beginning of the impact, we have been working to recover the remaining instances and volumes. A small number of remaining instances and volumes are hosted on hardware which was adversely affected by the loss of power. We continue to work to recover all affected instances and volumes and will be communicating to the remaining impacted customers via the Personal Health Dashboard. For immediate recovery, we recommend replacing any remaining affected instances or volumes if possible.
As far as I can tell, the number of a Mechanical / Hardware Failure is far higher than Software. And it is always Power, UPS, Generator, BBU, Raid Card Failure etc.
Why is it that we keep hearing failure in these segment? And it doesn't seems anything have been done? Are there any Innovation happening in this space?
> We’d like to give you some additional information about the service disruption that occurred in the Tokyo (AP-NORTHEAST-1) Region on August 23, 2019. Beginning at 12:36 PM JST, a small percentage of EC2 servers in a single Availability Zone in the Tokyo (AP-NORTHEAST-1) Region shut down due to overheating.
By the way, I'm pretty sure none of the actual AWS datacenters are in Reston proper. They are in Ashburn and other more sparse suburbs.
Source: I live in the DC area and regularly visit Reston and Herndon. There are large AWS offices in Herndon but not so many datacenters. Real estate in Reston is pretty expensive.
Pretty sure you are right. The physical AWS data centers I know of in Reston area are:
4 DCs on Smith Switch Rd in Ashburn, VA
2 DCs (IAD54 and IAD67) elsewhere in Ashburn, VA
3 DCs on West Severn Way in Sterling, VA
3 DCs on Dulles Summit Ct in Sterling, VA
2 DCs on Prologis Dr in Sterling, VA
2 DCs on Relocation Dr in Sterling, VA
2 DCs (IAD69 and IAD76) elsewhere in Sterling, VA
1 DC in South Riding, VA
4 DCs on Westfax Drive in Chantilly, VA
3 DCs on Mason King Ct in Manassas, VA
2 DCs elsewhere in Manassas, VA
(Not that data loss or downtime mattered for this instance, which was just used for internal testing.)
Power-up is very stressful for hard drives, so it's not too surprising that some failed when the power turned back on. EBS does offer spinning rust storage options, so maybe mLab was using those for some of those failed volumes. I don't know if the same is true for SSDs or not.
I mean the failure of backup is not a surprise. Entropy is just everywhere.
talk to some datacenter admins and you'll learn there's a lot more bailing wire and hope-for-the-best out there than you would think.
Same deal as those 50 point inspections your mechanic does: some things are easier to inspect than others, some people do a more thorough job than others, etc.
This guy sounds like if he'd self hosted he'd be complaining about an HDD failure. It happens- you need to design around it. Luckily, EBS volumes, snapshots and AZs make all of this pretty straight forward.
The EBS documentation states that there's an expected AFR of 1 to 2 per thousand volumes, so you should plan accordingly. Replicate any data that will cause harm to your business if lost to other sources. Keep backups.
That's the reality of the life in a data center. So yeah, either accept that stuff like this happens or build for stuff like this happening. Engineering around physical problems in the cloud environment is far easier than in the data center environment.
Anyone thinking this doesn’t happen with private data centers is either very green or selectively excusing problems.