http://fox61.com/2016/08/08/delta-airline-reports-outage-eve...
The outage began at 0230 and is presumably ongoing and beyond the capacity of their backup generators to cope with.
http://fox61.com/2016/08/08/delta-airline-reports-outage-eve...
The outage began at 0230 and is presumably ongoing and beyond the capacity of their backup generators to cope with.
Delta airlines, with $40B of revenue last year, has less than 8 hours of backup power for their mission critical computers?
Even on classic DBs, setting up a failover with a lag of 15 minutes isn't a herculean effort. It's expensive, yes, but probably worth it...
Plus sometimes backup power isn't reliable (e.g. generator problem) or they have different tiers of coverage (and Delta didn't pay for it). We'll just have to wait and see.
There are things that are fairly easy to make redundant. I don't know anything about their infrastructure, but I doubt building a failover would be trivial.
Even if you had it in place, the risk of something occuring as you switch back and forth might not be worth it if you expect your regular setup to come back up within a reasonable time frame. Most people around here know how close to impossible it is to replicate real world usage on a test system, and how much overhead it adds to any change you make to your production system if you want the others to follow.
I've worked at several places where there has been battery backup, diesel generators, and a a disaster recovery site with generally nice and thorough strategies for replication including really expensive "sysplex-mode" for mainframes, topped with regular contigency tests where one half of the site is shut down.
And somehow, at every place, eventually something bad happened in a way that took down the applications.
It could be a power outage in the exactly the wrong circuit, fluctuating power messed up exactly the wrong storage system, someone dug up exactly the wrong fiber, the fire extinguisher set off by mistake and sprayed just the wrong cabinet with water fog through a door(!) that was supposed to automatically close, the diesel generator caught fire and the backup battery had to be taken offline, etc etc.
I'm not saying that you shouldn't try, but it's damn hard.
One day during a power cut the generator kicked in as it was supposed to, but a minute later it spluttered to a halt. After much fiddling around by maintenance staff they gave up and we were all told to go home.
The postmortem revealed that the diesel had become contaminated with all sorts of gunk - water, grit/sand and lumps of unidentifiable "stuff". The best laid plans and all that.
Edit: And cigarette butts as well.
How often did you test that?
> has less than 8 hours of backup power for their mission critical computers?
It's plausible that they thought they had that, but some part of it failed.
Not parent, but here is a single data point. When I was a student, I worked for some sysadmins who managed a small, internal data center. ~200 IBM servers, a SAN, an old VAX, etc. It had a backup gas generator that was tested loudly, every single week. It was also tested whenever the power went out for too long.
Obviously lots can go wrong, even with the best plans, but at least at the tiny little place I worked, the generator was tested every week.
Mind you, I'm not saying they don't have one, just that it can be difficult to get -- there can also be limits on the number of hours you can run them.
> According to the flight captain of JFK-SLC this morning, a routine scheduled switch to the backup generator this morning at 2:30am caused a fire that destroyed both the backup and the primary. Firefighters took a while to extinguish the fire. Power is now back up and 400 out of the 500 servers rebooted, still waiting for the last 100 to have the whole system fully functional.
http://www.flyertalk.com/forum/27032000-post135.html
Update: Confirmed.
> But Georgia Power said the issue had to do with Delta's own equipment, not a larger power outage. "We believe Delta Air Lines experienced an equipment outage; other Georgia Power customers were not affected," said John Kraft, spokesperson for the utility. Georgia power has staff on site trying to assist Delta, he said.
https://www.washingtonpost.com/news/the-switch/wp/2016/08/08...
http://news.delta.com/730-am-et-update-outage-affects-depart...
A power outage in Atlanta, which began at approximately 2:30 a.m. ET, has impacted Delta computer systems and operations worldwide, resulting in flight delays
8:40 a.m. ET UPDATE: A Delta ground stop has been lifted and limited departures are resuming following a power outage in Atlanta that impacted Delta computer systems and operations worldwide. Cancellations and delays continue.
That being said, I expect them to upgrade their generators to handle outages like this, and distribute more of their IT infrastructure. Unlike the typical summer storms in the Atlanta area that delay flights, this affected them worldwide.
The "all over the world" bit is why there's a single point of failure.
A while back there was a story that made the rounds of aviation geeks, about Delta flying an empty 747 to Korea. That was to replace a 747 which had been badly damaged by hail, to the point that it would be unable to operate its scheduled return flight to the US.
Do you want to guess how many people, parts and places were involved in the "simple" task of dealing with this problem?
At first, the replacement 747 is in storage in a "boneyard" facility in Arizona, due to having been recently retired. So first it has to be pulled from the storage facility, put through basic airworthiness checks and fueled up, and then a flight crew has to be present to fly it to a Delta hub where it can be readied for a trans-Pacific flight.
The hub in question turned out to be Minneapolis. There, the plane has to undergo more work to get it ready for a long flight, and now multiple flight crews have to be present, since they need to rotate in and out over the duration of the flight (that's how you do long flights). Oh, and Minneapolis isn't normally a 747 base; it only gets them during peak travel seasons and on the occasional charter. So crews probably have to be brought in, stores and maintenance setups need to be brought online, etc.
Then the plane can -- finally -- fly out to replace its damaged counterpart, pick up any stranded passengers and bring them to the US. Which will mean flying into yet another hub, since the flight doesn't go back to Minneapolis.
Meanwhile the damaged plane is still sitting there in Korea, and needs to be repaired on-site to get it into minimum airworthy condition to fly home (empty of passengers). It's going to need parts, maintenance crew, flight crews, etc. just like the replacement plane did.
And the deeper you dig the more stuff you'll find like this. Running an airline with global, or even national-across-the-US, service is not something you can decentralize to avoid problems at one operations center. The amount of coordination just of people, parts and planes across widely disparate locations requires centralized operational control instead of devolved regional centers with high autonomy.
At this point I consider a company as large as this having such a rudimentary single point of failure to be incompetence in the IT department. We wouldn't be so forgiving if delta needlessly kept all of its pilots in one city during the night so a single storm wiped out every flight.
You're still out of luck when the centralized admin center goes down, though. That's the place that is the source of all the humans performing the coordination and dispatching work. Having a bunch of extra data centers and backup generators around the country will not cause those humans to become accessible.
And building out full redundant continuity of everything, including the humans, is not something that tends to happen outside of major governments.
Also, "centralize administration" just means that you can control everything from a single location. It doesn't preclude being able to control from multiple locations.
Think of AWS, you can control everything across multiple data centers from a centralized interface from anywhere with an Internet connection, even if entire data centers go down.
A sane system should essentially allow delta to operate from many possible locations seamlessly as long as they have the human operators required.
The generators have been know to fail hard, and then the entire DC is down. On top of that, generators often are only able to run for a limited amount of time hours to days [2].
[1] https://en.wikipedia.org/wiki/Flywheel_energy_storage [2] http://www.graniteblockglobaldatacenter.com/powersystems.htm...