No. Not even rarely. If they lost hardware because of this something much different than just loss of power happened on their servers' mains rails.
No. Not even rarely. If they lost hardware because of this something much different than just loss of power happened on their servers' mains rails.
None of those things are fool proof and in case of a large scale outage, especially one that lasts more than a couple of minutes there is a fair chance that you will find the limitations of some of your contingency plans.
AWS has a lot of very skilled professionals who are quite familiar with those issues so I'd be quite surprised if it turned out to be something that simple but you always want to have a contingency plan you've tested for what happens if core infrastructure like that fails and takes multiple days to recover.
1. One nasty example: fuel issues which aren't immediately obvious so someone doing a 5 minute test wouldn't learn that the system wasn't going to handle an outage longer than that.
Most notably we had a large NetApp array that would not boot, once we replaced the controllers we had lost 13 out of the 60 hard drives in the array. They would no longer spin up. Like physically seized. Because they had been spinning for so long, they would have likely kept spinning just fine, but with power gone, it was done.
Fans are another fun one, anything with ball bearings really. Power supplies are another issue, due to the sudden large inrush of power when the switch was flipped back to on, some power supplies had their capacitors go up in smoke.
This is not rare, this a common occurrence. When you have an absolutely massive footprint with thousands upon thousands of servers, that have been running for a long time, there will be things that just don't come back once they have stopped running or when the electricity is gone.
At AWS scale it’s highly likely there are more than a few of these.
It's really common to find thyristors in devices to limit in-rush current, even in cheap electronics. You can bet that PSUs in data center equipment use them.
Yes, absolutely.
The post you’re responding to never implied that it happens to _most_ of them, or even many of them, just that when you have a gigantic farm, it’s not unreasonable to see a small handful of hardware instances release up their magic smoke when they’re all coming back from being powered down.
MTBF is a thing.
If electronics (not just computer) is going to fail its almost always at power on or power off.
And since DC is one big electronics warehouse, stuff breaks. Old spinning drives were most prone to this, with power supply and motherboards being second.
Due to me not having that much free space and also enjoying a good night's sleep, i have cron set up to put the servers offline (rtcwake) while i'm sleeping, so that they wouldn't just go on droning in the background and to start up in the morning.
Since they do this every day with no exceptions, that's a decent amount of power cycling, in addition to them working all the time otherwise, which has been going on for about 5 years or so.
So essentially each of them has been on for about 30'000 - 35'000 hours and has power cycled around 1800 - 3600 times, because for the longest time i've had a cron restart set up right after the rtcwake sleep period finishes, due to some odd bug in the kernel that otherwise caused their CPUs to get stuck at 100% usage after waking up.
Over these 5 or so years, i've had:
- 2 PSUs die (possibly one due to a power surge though, before i had UPSes in front of them), seemed to work fine but then one morning didn't start up
- 2 HDDs die (Seagate Barracuda 1 TBs, each box has 2-3 of them), typically this isn't abrupt, but the HDDs lose power when trying to spin up and loop like that for a while before eventually deciding to work, before eventually not being able to do that anymore at all; oddly enough, Clonezilla could still mirror such a disk after a few attempts
- 1 motherboard die (not sure precisely why, then again, for a while in my university i had the servers (without cases) up next to a slighty ajar window, even in winter as well as in summer, while they were otherwise only passively cooled (x86, but 35W TDP 200 GEs)
- some problems with bearings, once in a cheap CPU fan before i had passive cooling and then again in one of the current PSU fans, sometimes when starting up it sounds like a miniature chainsaw (checked the fan, isn't actually seized up or anything, just cheap bearings)
- some problems with the power button (well, more like the pins, since i didn't have a proper button at one point, just a screwdriver to close the circuit with), though i solved that by closing the circuit on the ATX connector itself by putting a little wire in there
All of the failures have happened around power cycling, never in the middle of otherwise normal operation.To my surprise, nothing has actually melted or caught fire, the last time when anything close to that happened was ages ago with my 9800 GT GPU which for some reason started getting really bad thermals (90 degrees C) and so i had to reflow solder it in the oven a few times, before eventually it just refused to work and artifacted out on boot. Or maybe i could mention a MOLEX connector stopping working, but i might just have a splitter in there due to not enough SATA connections from the PSU, something that people typically advise against.
So, to sum all of that up, yes, consumer hardware will definitely die from regular use under normal circumstances, probably moreso with regular power cycling of parts that aren't really made for that (especially the regular HDDs, which seemed to die right after restarts). Also, my case might be an outlier (and a sample size of 1), due to me not caring that much in university about keeping the servers in a good environment, but i bet that most consumer hardware has to deal with sub-optimal conditions in one way or another - accumulation of dust, bad airflow, air humidity, ambient temperature and power grid fluctuations all should also be taken into account.
I actually had to stagger the startup times for my servers, since otherwise they'd max out their power draw at the same time in the morning (e.g. when launching Docker containers and spinning up VMs), which also lead to some weirdness. Of course, the quality of the parts themselves probably also matters (like a reputable brand PSU vs the cheaper options that i could afford) as much as how they're used (prolonged periods of full load while everything starts up), as do other factors, like vibrations of the HDDs, for example.
Also:
> Did you think this over before hitting the reply button?
This is ad hominem, and we should avoid it here.
Recently we has a violent power outage where I work due to the tornados in the midwest. A few of our facilities have industrial hardware with batteries that will last for 72 hours and after that point it’s a world of unknowns what will happen. The batteries died after a lot of effort to try and restore power, but everything came up just fine.
However, a random Cisco switch elsewhere in the building had at least one power supply fail.
This was in one facility where most of these have been running for 4+ years straight (12 years for the system with a 72 hour battery.)
I find it hard to imagine that in the scale of even one AZ that it wouldn’t be possible that at least one system has this happen.
You have outed yourself as someone with zero experience operating computers, either individually or at scale. Computers have moving parts, whether fans or hard disk drives. Just because those were moving doesn't mean they will begin moving again from a dead stop. There are also things like dead cmos/nvram batteries that can prevent machines from automatically starting when power is applied.
This is hostile enough that I'd have trouble squaring it with the site guidelines.
It's especially bad because, as others have been saying, your belief isn't supported by real lived experience for many of us. A data center has enough devices that even a low probability error rate will fairly reliably happen, and that doesn't usually hurt the manufacturer's reputation because they never promised 100% and will replace it under warranty. I've seen this with hard drives especially but also things like power supplies and on-board batteries, and even things like motherboards where a cooling/heating cycle was enough to unseat RAM or cause hairline fractures in solder traces.
This can be especially bad with unplanned power outages if the power doesn't instantly go out and stay out for the duration. Especially for hard drives hitting the power up/down cycle a few times was a good way to have extra failures.
Unscheduled power outages can be even worse
Other hardware is similar.
But if I lost power to thousands of servers, I'd expect some number of them to fail. I've even lost servers when losing power to a single rack.