AWS power failure in US-EAST-1 region killed some hardware and instances
theregister.com
theregister.com
My 2 cents. Outages happen. Network glitches happen. Bad configs and bad updates happen. However, power issues should not really happen. One of the primary cost saving areas of going to the cloud is not having to do on-prem power, such as UPS, generators, maintenance etc. Not having to do on-prem cooling is another thing. These should be solved from the customer’s perspective when going into a professional data center and are the things you don’t want to worry about anymore.
I used to work at AWS. The most worked-up I ever saw Charlie Bell (the de facto head of AWS engineering) was in the weekly operations meeting, going over a postmortem which described a near-miss as several generators failed to start properly in a power interruption. In this meeting - with nearly 1,000 participants! - Charlie got increasingly irate about the fact that this had happened. For the next few weeks, details of some sort of flushing subsystem for backup generators was an every-time topic.
Sadly, Charlie left AWS and now works at Microsoft. I can't help but wonder what that operations meeting is like now without him.
This suggest to me DataCenter, as an infrastructure and building on itself still have plenty of room for improvements.
Leading up to Y2K, I remember concerns about spinning hard disks not being able to start up again.
And if the power is flaky with spikes and brown-outs, I understand that's a problem.
But is either of those relevant to AWS?
* increased failure rates of drives (or other hardware) due to thermal cycling
* cache power failures leading to corrupted data
* previously unknown failures in recovery processes (like firmware bugs that might be described as a "hardware problem")
* cooling failures leading to hardware failure
Most of these are pretty rare issues, but it becomes more probable to happen at least once when you're running freekin' us-east-1
I remember IBM was renting private planes to get us replacement hardware from around the country (we were a very large customer).
TLDR if electronics are near-death, a power cycle is likely to kick it over the edge.
Don't know the truth to that, but it seemed like a reasonable explanation at the time.
I have an old fan that will need replacing one day that takes about 15 minutes to reach full speed as the bearings are gone. Everytime I turn it off I wonder if the next time I turn it on it will be time to replace it.
Further to that thought, there must be hardware in that centre which is failing and being replaced all the time due to other conditions, we just won't hear about it because it's part of normal maintenance.
After letting devices (especially that rarely do this) cool, and then bringing them back up to temperature some percentage of components may fail or operate out of spec.
And mechanical parts have similar issue. A motor that's been happily running for months/years may not restart after being stopped.
It's also possible that a power fault itself could damage hardware, although in my experience this type of issue is far less common.
Here's a few references at random:
https://vbn.aau.dk/ws/portalfiles/portal/108034396/ESREF2013...
https://www.infineon.com/dgdl/s30p5.pdf?fileId=5546d46253360...
https://www.sciencedirect.com/science/article/pii/S003811010...
[1] Yes yes always have backups, and RAID, and replication something about not being able to count that low. Betcha most'a'ya don't have any of that on your MBP though except for a maybe a backup that lags hours behind your work.
Suppose that the "slamming on the brakes" here is that the power lines connecting a data center to the grid get physically cut. I can see how that would cause a power spike on the grid side of the cut, but would it really create a spike on the datacenter side of the cut?
Or assuming that the wires in a cable aren't cut exactly simultaneously, does it depend on the particular phases of the AC power at the time of each wire in the cable being cut?
The motors generate appreciable back EMF when suddenly unpowered. Likewise an SMPS is an inductive load. Spikes from power supplies going off is probably what killed the disks, assuming they didn't just decide not to power on. Some disks will run for 10 years but power them off once and they don't come back up.
AWS RDS in multi-AZ deployment gives you two availability zones. Aurora gives you three. What kind of scenario would be used to justify three AZ’s for the purposes of high availability?
If you can't handle a single az failure there is no way you are going to handle failing over across different cloud providers correctly.
Not advocating for multi-cloud though...
Given the number of AWS global services that have dependencies on infra in US-EAST-1 (and, from the impacts of this and other past outages, seen vulnerable to single-AZ failures in US-EAST-1) that's...less avoidable for certain regions/AZs than one might naively expect. Most clouds seem to have at least some degree of this kind of vulnerability.
This is not true. Amazon is not being upfront about what happened here. It was simply not a single AZ failure. Our us-east-1 ELB load balancers were hosed and were unable to direct traffic to other AZs - they simply stopped working an were dropping traffic. We tried creating load balancers in different AZs and that didn't work either.
How can you be resilient to single AZ failures if load balancers stop working region wide during a single AZ outage?
By definition, you would have to either go lowest-common-denominator, or build complicated facades in front of like services.
If you're going lowest-common-denominator, then multi-old-school-hosting would be far cheaper.
vp: no.
A few hours downtime is not going to justify double cost, and worse whose benefits can only be demonstrated during those few hours.
It will get shutdown immediately.
And what is more important, corporate world really don't care that much about downtime, not as much as they care about who to assign the blame, in this case AWS is a perfect irresistible externality much like a natural disaster.
Disclaimer: Ex-AWS employee
IMO, most companies just aren't sensitive enough to downtime that multi-AZ + multi-region deployment within a single cloud provider isn't good enough.
No. Not even rarely. If they lost hardware because of this something much different than just loss of power happened on their servers' mains rails.
At AWS scale it’s highly likely there are more than a few of these.
It's really common to find thyristors in devices to limit in-rush current, even in cheap electronics. You can bet that PSUs in data center equipment use them.
Yes, absolutely.
The post you’re responding to never implied that it happens to _most_ of them, or even many of them, just that when you have a gigantic farm, it’s not unreasonable to see a small handful of hardware instances release up their magic smoke when they’re all coming back from being powered down.
MTBF is a thing.
If electronics (not just computer) is going to fail its almost always at power on or power off.
And since DC is one big electronics warehouse, stuff breaks. Old spinning drives were most prone to this, with power supply and motherboards being second.
Due to me not having that much free space and also enjoying a good night's sleep, i have cron set up to put the servers offline (rtcwake) while i'm sleeping, so that they wouldn't just go on droning in the background and to start up in the morning.
Since they do this every day with no exceptions, that's a decent amount of power cycling, in addition to them working all the time otherwise, which has been going on for about 5 years or so.
So essentially each of them has been on for about 30'000 - 35'000 hours and has power cycled around 1800 - 3600 times, because for the longest time i've had a cron restart set up right after the rtcwake sleep period finishes, due to some odd bug in the kernel that otherwise caused their CPUs to get stuck at 100% usage after waking up.
Over these 5 or so years, i've had:
- 2 PSUs die (possibly one due to a power surge though, before i had UPSes in front of them), seemed to work fine but then one morning didn't start up
- 2 HDDs die (Seagate Barracuda 1 TBs, each box has 2-3 of them), typically this isn't abrupt, but the HDDs lose power when trying to spin up and loop like that for a while before eventually deciding to work, before eventually not being able to do that anymore at all; oddly enough, Clonezilla could still mirror such a disk after a few attempts
- 1 motherboard die (not sure precisely why, then again, for a while in my university i had the servers (without cases) up next to a slighty ajar window, even in winter as well as in summer, while they were otherwise only passively cooled (x86, but 35W TDP 200 GEs)
- some problems with bearings, once in a cheap CPU fan before i had passive cooling and then again in one of the current PSU fans, sometimes when starting up it sounds like a miniature chainsaw (checked the fan, isn't actually seized up or anything, just cheap bearings)
- some problems with the power button (well, more like the pins, since i didn't have a proper button at one point, just a screwdriver to close the circuit with), though i solved that by closing the circuit on the ATX connector itself by putting a little wire in there
All of the failures have happened around power cycling, never in the middle of otherwise normal operation.To my surprise, nothing has actually melted or caught fire, the last time when anything close to that happened was ages ago with my 9800 GT GPU which for some reason started getting really bad thermals (90 degrees C) and so i had to reflow solder it in the oven a few times, before eventually it just refused to work and artifacted out on boot. Or maybe i could mention a MOLEX connector stopping working, but i might just have a splitter in there due to not enough SATA connections from the PSU, something that people typically advise against.
So, to sum all of that up, yes, consumer hardware will definitely die from regular use under normal circumstances, probably moreso with regular power cycling of parts that aren't really made for that (especially the regular HDDs, which seemed to die right after restarts). Also, my case might be an outlier (and a sample size of 1), due to me not caring that much in university about keeping the servers in a good environment, but i bet that most consumer hardware has to deal with sub-optimal conditions in one way or another - accumulation of dust, bad airflow, air humidity, ambient temperature and power grid fluctuations all should also be taken into account.
I actually had to stagger the startup times for my servers, since otherwise they'd max out their power draw at the same time in the morning (e.g. when launching Docker containers and spinning up VMs), which also lead to some weirdness. Of course, the quality of the parts themselves probably also matters (like a reputable brand PSU vs the cheaper options that i could afford) as much as how they're used (prolonged periods of full load while everything starts up), as do other factors, like vibrations of the HDDs, for example.
Also:
> Did you think this over before hitting the reply button?
This is ad hominem, and we should avoid it here.
Recently we has a violent power outage where I work due to the tornados in the midwest. A few of our facilities have industrial hardware with batteries that will last for 72 hours and after that point it’s a world of unknowns what will happen. The batteries died after a lot of effort to try and restore power, but everything came up just fine.
However, a random Cisco switch elsewhere in the building had at least one power supply fail.
This was in one facility where most of these have been running for 4+ years straight (12 years for the system with a 72 hour battery.)
I find it hard to imagine that in the scale of even one AZ that it wouldn’t be possible that at least one system has this happen.
You have outed yourself as someone with zero experience operating computers, either individually or at scale. Computers have moving parts, whether fans or hard disk drives. Just because those were moving doesn't mean they will begin moving again from a dead stop. There are also things like dead cmos/nvram batteries that can prevent machines from automatically starting when power is applied.
This is hostile enough that I'd have trouble squaring it with the site guidelines.
It's especially bad because, as others have been saying, your belief isn't supported by real lived experience for many of us. A data center has enough devices that even a low probability error rate will fairly reliably happen, and that doesn't usually hurt the manufacturer's reputation because they never promised 100% and will replace it under warranty. I've seen this with hard drives especially but also things like power supplies and on-board batteries, and even things like motherboards where a cooling/heating cycle was enough to unseat RAM or cause hairline fractures in solder traces.
This can be especially bad with unplanned power outages if the power doesn't instantly go out and stay out for the duration. Especially for hard drives hitting the power up/down cycle a few times was a good way to have extra failures.
Unscheduled power outages can be even worse
None of those things are fool proof and in case of a large scale outage, especially one that lasts more than a couple of minutes there is a fair chance that you will find the limitations of some of your contingency plans.
AWS has a lot of very skilled professionals who are quite familiar with those issues so I'd be quite surprised if it turned out to be something that simple but you always want to have a contingency plan you've tested for what happens if core infrastructure like that fails and takes multiple days to recover.
1. One nasty example: fuel issues which aren't immediately obvious so someone doing a 5 minute test wouldn't learn that the system wasn't going to handle an outage longer than that.
Most notably we had a large NetApp array that would not boot, once we replaced the controllers we had lost 13 out of the 60 hard drives in the array. They would no longer spin up. Like physically seized. Because they had been spinning for so long, they would have likely kept spinning just fine, but with power gone, it was done.
Fans are another fun one, anything with ball bearings really. Power supplies are another issue, due to the sudden large inrush of power when the switch was flipped back to on, some power supplies had their capacitors go up in smoke.
This is not rare, this a common occurrence. When you have an absolutely massive footprint with thousands upon thousands of servers, that have been running for a long time, there will be things that just don't come back once they have stopped running or when the electricity is gone.
Other hardware is similar.
But if I lost power to thousands of servers, I'd expect some number of them to fail. I've even lost servers when losing power to a single rack.