Delta Computer Glitches Force Flight Halts Third Year in a Row
bloomberg.com
bloomberg.com
> Still, Delta isn’t the only U.S. carrier to have suffered from technical glitches. In December 2017, a fire at Atlanta’s airport, the world’s busiest hub, caused a major electrical disruption, crippling services and stranding thousands of passengers of Delta as well as rivals including Southwest Airlines Co.
It seems to me that Southwest's systems being brought down by fire and electrical disruption is a different "technical glitch" than the kind of systems problem Delta just had. A better comparison would be July 2016, when a router failure created a "chokepoint that crippled hundreds of the company's software applications":
https://www.dallasnews.com/business/southwest-airlines/2016/...
To some degree, sure. But they really should be capable of a straight failover (forced, if necessary) to geographically separate components.
AFAIK the failure at ATL had nothing to do with the ability for the IT network to failover. It was because the power at the airport was off, and therefor the airport was unable to fly planes. And when the world's largest airport goes dark, it causes ripples in delays across the entire country because of misplaced planes/staff/passengers. And you can't just "failover" hundreds of planes and thousands of passengers to a different geographic location.
We're really talking about two completely different types of failure here, and the other comment is right in that it's weird to compare the ATL fire to last night's failure.
Even for non-aviation-or-tech-knowledgeable people, I think most still care who is at fault for such disruptions. The article frames the ATL fire as Delta/Southwest's fault. It puts the ATL fire in the same category as other, airline-specific issues that were caused by Delta's IT systems. But that just isn't the case. The ATL fire affected every airline at ATL (but for some reason the article only calls out SW and DL), and was the fault of the airport, not the airlines. It's more similar to bad weather shutting down an airport than it is to the IT issues that Delta faced last night.
It just feels like bad reporting to include that comparison.
I had the weirdest job interview for a contracting job at Delta. I passed two technical phone interviews and was scheduled for an in-person interview. The night before it dumped a ton of snow. I tried to cancel, but didn't have the manager's contact info. I left an hour early to get there, and still I was 5 minutes late.
They let me sit in the lobby for another twenty minutes, then the Manager's "administrative assistant" came and fetched me. The Manager then asked me some of the most condescending questions, like "what operating system do you use? Have you ever used Linux?" He then finished with a snide comment and that our time was up and he had a meeting to go to.
I called the recruiter from my car. They decided pass, "because the Manager said I didn't apologize enough for being late." That was the last I heard from the recruiting company. They ghosted me. Wouldn't talk to me or respond to my email. It was the most unprofessional situation I've run into in 25 years of contract/consulting. Just plain weird.
Note, this job was not to code, but write documentation about the code. They wanted a C++ programmer to document the existing system that schedules and reroutes flights during irregular ops, that was written by some guys from Bell Labs running on some fault-tolerant hardware. He had sold the management on converting everything into Java running on commodity/cloud servers.
Every time I see their system meltdown... I'm reminded of this.
https://www.linkedin.com/pulse/people-ghosting-work-its-driv...
The real showstopper is if the Operations Center can't deliver the dispatch release; that's the legal document prepared by the airline dispatcher that grants the flight permission to initiate the flight. Without that, it's a no-go. Even if Deltamatic were down, Flight Ops could fax the releases over, as the dispatch planning system is merely tied into Deltamatic to make local printout easier; it doesn't actually run on Deltamatic.
These days with e-ticketing and more stringent DOT/Customs standards on accurate pax manifests, etc, having Deltamatic (ticketing/res) down is a show-stopper, but it wasn't always that way and the way these articles are written it leaves, in my mind, the open question about what parts were actually fubared.
(me: ex-aircraft dispatcher for a DL-owned regional carrier.)
The impression I get from these repeated fiascos is the whole industry still runs on mainframes and java applets held together with duct tape and bubblegum, and that their executives are dinosaurs with no appreciation for what it takes to build robust failure-tolerant systems. But maybe it's just a skew in the reporting.
In the world outside Silicon Valley, and outside of Venture Capital-funded projects, the type of code and infrastructure everyone here idolizes is the exception by a wide margin.
* Sharding and clustering.
* Queuing between services.
* Geographic distribution across multiple datacenters.
* Continuous restoration from backups.
* Unit testing.
* Source control.
* ...
Those practices are not language or system specific, but they may not be easy to follow when using older technologies.
You’re right though. That said, these systems are crufty. It’s true that they get the job done often, and updating them doesn’t inherently mean that they’re better, but they fail a lot. One recent example; I had a reservation cancelled 24 hours after booking three times, with no explanation why. “Normally there’s a note here explaining... but this time, nothing. This just happens and we don’t know why.”
The person telling you doesn't know why, and there is no way for them to directly check, but there are a variety of reasons reservations get canceled, reasons they airline doesn't want to pass on to you. Top of the list is being pumped for an airline employee who needs to be somewhere. Then come premium customers such as government agencies/police/military. Then rich people (premium travelers/club members etc). When such people get a flight at the last minute, they are bumping someone else. That someone isn't told why.
In this case, I don’t think that’s what happened; I was able to re-book the exact same flight all three times. In those cases, I should at least be notified my flight was cancelled, but was not. The third rep suggested it was a KLM/Delta integration issue. This is also only one of many issues I’ve run into over my years of pretty extensive travel.
Regardless, you are right; it could be a white lie.
It's a risk aversion issue. Large software projects are horrendously difficult to get right, and managing airlines is a horrifically complicated kind of project. Fault tolerance aside, just correctly modeling the problem is a mess and by the time you have a correct solution, the problem itself may well have changed. So you end up having to throw a lot of resources at a very complex problem and know that you may entirely fail at it, which is not the kind of gamble execs love.
Banks for a long time thought the same. Then banking startups with code written on a clean slate in modern architectures emerged and are eating the bank's lunches as their infrastructure (more often than not a mix of mainframe, weird custom middleware and half-modern webstacks) cannot keep up in development pace. For example, in Germany Commerzbank (once a powerhouse) got kicked out of the DAX with former fintech startup now giant Wirecard replacing them.
Airlines are only "kept safe" from modernization because the cost of entry is so unbelievably huge - to start a bank one needs only a couple of employees plus a couple million euros founding money, but to start an airline one needs vastly bigger sums.
I think that's one of the keys to their success. Large banks can't just say, "well, we're not going to be supporting A, B and C so we can do a total rewrite."
My info is maybe 10years old, but I doubt much has changed.
These infrastructure type things are incredibly complicated. Not in the “Newfangled” framework way, but rather the business implications of real world impact of potentially bad code.
I consider it one of the highlights of my career going thru the Delta Northwest Merger without shit hitting the fan. We were writing decrementer & while loops instead of for loops because there was a measurable performance & stability impact.
There’s something to be said about software that has withstood the test of time.
I’m impressed that the team is able to handle outages as well as they do. Having had some experience working with the folks at Delta & Northwest, I’d say they’re pretty impressive people. I’d be happy to steal some for my startup if they’d be willing to part with the flight benefits & tenure.
Full Disclosure: I was not a Delta Employee, but worked for a startup that built the checkin systems (which got acquired by NCR). I got to work with Delta as a client.
No idea what their backend looks like though.
The problem is that those systems are pretty segmented from the systems that schedule/dispatch planes and staff. Those systems are incredibly complicated due to the complex nature of managing the status/location of 1,000+ airplanes, millions of pieces of cargo, thousands of flight attendants and pilots, etc. And those systems also often have to interface with antiquated airport or FAA systems. As a result, making changes to the reservations or flight management system is a lot different than working on a web app.
With SABRE up and running, IBM offered
its expertise to other airlines, and
soon developed Deltamatic for Delta Air
Lines on the IBM 7074, and PANAMAC for
Pan American World Airways using an IBM
7080.
https://en.wikipedia.org/wiki/Sabre_(computer_system)https://www.linkedin.com/pulse/mainframes-problem-solution-j...
That wouldn't make that server competitor neither for IBM Parallel Sysplex, nor even for Oracle Exadata.
> A number of financial institutions have discovered and acknowledged this fact and are now taking the first steps to address this situation by making the decision to go back to z/TPF, z/OS and UNIX because of their bulletproof reliability and transaction processing capabilities. This to the chagrin of the bulk of their application development staffs who are bolting out of these organizations for the more sexy technology environments. Yes, IBM Assembler, COBOL and other antique languages are making a resurgence at these brave organizations as they supplant and even replace Java, .NET, Python and other development environments.
Back to the good old days of IBM Assembler?!
Are they ditching source control and unit testing too?
When I left that company a few years after that, there was still no source control for the RPG or COBOL code.
"Using Visual COBOL in Modern Application Development" https://www.microfocus.com/documentation/visual-cobol/visual...
High Level Assembler to be correct. But kidding aside, the slide 40-41 is interesting http://www.ibm.com/software/htp/tpf/tpfug/tgf18/TPFUG_2018_M...
Interestingly, there was a very, very long line at the ticketing counter of Delta this morning. I asked a few passersby if they were put up at a hotel like me. Every person I asked was in the same situation.
There's probably the odd flight that is held back if there's a large number of passengers that will miss it due to a late connection--and there's no later flights with seats available. I would imagine the cost of holding the 11pm LAX-SYD flight because 20 people are 10 minutes late on the NYC-LAX connection is a lot less that paying for hotels, especially when the next flight may not have 20 seats available anyway. Though then again, most carriers have gate pressure at LAX and they might not be able to hold the flight anyway.
It's an interesting optimization problem to solve!