United Airlines System-Wide Computer Problem
flyertalk.com
flyertalk.com
Went like this: Guy who shows around the new datacenter/ops guy demonstrated how the emergency power off works by lifting the protection plate. Protection plate unhinges suddenly and droppes onto emergency power off button. Hilarity ensues.
Confused the heck out of us when we were trying to figure out why some of our servers went on UPS power randomly at night. Turns out we'd get the notifications on nights the cleaning crew decided to flick the lights off then on again.
Good thing we deployed those EMCs in pairs. But then curiosity gets the better of a DC ops guy wondering how the plate got jammed. So he punches the plate on the 2nd EMC causing it to similarly jam and power-off. Doh.
When we told EMC about this flaw in the design, they deployed a fix to the production line - a rubber stopper under the plate next to the switch to protect the switch from the protection plate.
There's a plate on the switch now, praise Eris.
Had another instance where a pipe ran from the floor above right over the UPS. Construction workers on the first floor decided to poor some unusued paint into it (figured it was a drain pipe) and of course, we lost power again.
(yeah, I know security is already a huge deal, but as we come to trust software systems more and more, the safety/reliability factor will come more into play)
EDIT: This is also part of the reason I've been learning Elixir (http://elixir-lang.org/) since it's based on the highly-resilient Erlang and is designed to embrace failure. This was also informed by me reading Nassim Taleb's book "Antifragile" as well as "Thinking in Systems: A Primer" by the (late) Donella Meadows.
I actually don't know about this stuff, so any correction of my thoughts is appreciated.
That seems like a naive "MARKET WILL FIX IT" approach.
More likely, if the market does fix it, it'll be by having insurance companies deploy actual inspectors who know what they are doing and what sorts of problems to work for.
They might even be a fun combination of physical pentester/irl chaos monkey. Doesn't that sound like a fun job?
System audits would have to be standardised. There's bits of this in ISO9001, PCI compliance, FIPS, and so on. But the technology changes rapidly and the insurance companies don't have the expertise.
Doesn't come cheap though.
This actually applies to all businesses; speed improves, market fit improves (quality is just "what is good and valuable" after all), employee happiness improves (exponential gains from that alone). The effects are systemically positive.
But, no one cares, and the belief that we have to crack the whip to get people to make things faster, and we should reward the good people and punish the bad ones—will continue on as pure religious fallacy resulting in the failure or constant-operation-at-the-edge-of-failure of everything involving more than five people, probably until the end of the human race.
That's quite a wide gulf.
Would love to hear from people who have worked with developing / managing relatively high quality software with relatively modern methods.
Edit: Maybe NASA's Faster Better Cheaper comes to mind...
Deming -- W. Edwards Deming, that is -- is somewhere in the middle, around the right balance. He advocated for spreading a philosophy of systems thinking, scientific method, and statistical understanding, while simultaneously empowering employees by recognizing the power and responsibility of management and leadership, and understanding the motivation from a scientifically accurate psychological viewpoint.
It's a correct framework, and it's all aimed at driving quality by improving the things that directly impact it at a base level: fundamentally, how people work together, how they build systems that work, and how they're motivated (and demotivated) in reality.
Was the 9th bit special? Or just a standard bit in the byte?
If I'm not mistaken, some old consoles used a 9-bit RGB encoding, so this could theoretically help there. Minecraft also uses a 9-bit system for Redstone, afaik.
36 bits got us the 6-bit character (10 digits, 26 letters, and punctuation) with six characters in a word. Because of that, some OSes had six-character file names.
If you want to get upper- and lowercase, you need more than 6 bits. 9 is the smallest divisor larger than 6 of 36, so nine-bit characters made sense.
On such systems, file names could still use 6-bit characters, while applications used 9-bit ones. Also, some instructions could work on words, half words, quarter words, or sixth words.
One of my friends used to intern for a very large company that maintained software for flight control towers. His entire summer was spent writing bash tests for these old Fortran apps that kept planes from running into each other. Most of the mainframes still had tape.
That's the code that keeps track of our planes.
I think there is / will be a lot of money to be made trying to solve the problem of software security and reliability. This is obviously an extremely difficult problem, however the number of ancient systems that we currently have interconnected I think more large scale outages like this are inevitable.
Case in point, take this HackerNews fan favorite[1] about a school district using an "ancient system" and avoiding $M's in incremental costs.
[1] https://news.ycombinator.com/item?id=9705830
edit: As mentioned in [1], I assume that at least someone is aware of the cost-benefit of potential projects; in turn, I assume that someone would've pulled the trigger if the $'s make sense.
Currently says:
ATCSCC ADVZY 027 DCC 07/08/15 UAL GROUND STOP REVISION
DESTINATION AIRPORT: ALL AIRPORTS
FACILITIES INCLUDED: ALL FACILITIES
GROUND STOP PERIOD: 08/1200Z - 08/1315Z
REASON: USER REQUEST AUTOMATION ISSUESI have a sneaking suspicion that booking systems for most airlines run atop legacyware. It just seems like the type of thing that would've been put in place long ago and then be very expensive to migrate/updgrade.
http://www.pacbiztimes.com/2012/04/06/united-takes-a-step-ba...
Well, that answers that.
Meanwhile news articles and twitter complaints abound. http://mashable.com/2015/07/08/united-computer-problems-flig...
That was the worst of it, but almost every flight I saw on the way (both ones I was on and other flights at nearby gates) was delayed or overbooked or otherwise messed up in some way.
1. Do nothing. You're stuck here. We'll get you on the next flight. Oh, the next available flight isn't for over a week. Sorry. No, we can't pay for your rental car.
2. Upon renting my own car, and writing to customer service to complain I got a $125 voucher for United. Great. Not enough to buy an entire ticket, so it just ensures I'll have to continue giving money to this airline that failed to deliver what I gave them money for in the first place.
3. After several weeks of emails back-and-forth with customer service finally a manager agrees to issue a reimbursement for the rental car as a one-time exception to their policy of never doing that. (but not the $7 worth of gas. I guess they just have to draw a line somewhere? Oh, and they also rescinded the voucher. No big deal since I'm not too keen to fly with them again anyway)
So, yeah. Sometimes shit goes wrong when you travel. Everybody knows that. How airlines act to fix it makes all the difference. If you hamstring all your customer service reps so they can't actually solve someone's problem, it makes something that's already annoying way more frustrating.
Obviously, with travel sometimes things go wrong, but the quantity and severity of things going wrong at United makes me think there's something off about the way they're running their airline. For instance, the two times this year that they've had to ground all flights.
I have 1k status with United -- I fly internationally with them almost monthly. The problems I've had with United have been when bags have to be interlined to Brussels Airlines and occasionally (surprisingly) Lufthansa. I've also had Lufthansa somehow think it was a good idea for 2 toddlers to sit in scattered seats rows away from their parents (who were also rows apart as well, despite having checked in almost 24 hours before the flight and being Star Alliance Gold.) I've had Brussels Airlines say they were going to "gate check" a stroller only to have it show up days later. I've been stranded in Detroit back in the Northwest Airlines days when aircrews hadn't showed up to work. I've been stuck in Paris when the Air France pilots decide that salaries up to $300,000 per year just aren't enough. On a recent United trip from Hartford to Marseille, I was stuck in Hartford for % extra hours for an airplane that was stuck in Newark (just a <50 minute flight away.) I then missed a cascade of connections leaving me rather miserable. However, United sorted the problem and got me on my way as quickly as possible. Let's not forget Jet Blues antics on multiple occasions a few years ago: a 10.5 hour tarmac delay, a 7 hour tarmac delay among several other extremely long tarmac delays. AA had 14 long tarmac (over 3 hours) delays in February. United had zero long tarmac delays during the same period. Envoy/American Eagle was in last place for on-time arrivals last year.
I'm not defending United. I'm not disparaging the others. The fact is that the air transport industry is extremely complex and perceptions of quality are as varied as their are passengers in the sky.
Every airline sucks and every airline is great. Pick a day, pick a destination and roll the dice. When you fly often enough it seems like it all averages out to just one level of melancholic service; unless you're flying on Singapore Airlines -- then it just becomes sublime.
I fly out of UA hubs frequently and have had nothing but excellent service from them this year (over 50 segments flown this year, ~60K miles).
Definitely helps to have status too...
Personally, I consider SW to the be shittiest of them all. I hate having to fight for a seat...
Lastly, use google.com/flights by far my favorite booking tool now.
Edit: And its Twitter account has been relatively inactive, with more than 30 minutes since the last reply-to or general tweet...presumably a lot complaining tweets have come in in the last 30 minutes https://twitter.com/united/with_replies
"Departing DEN; taxied and then returned to gate. Pilot says nationwide failure of "three or four" computer systems. Only information from airport staff is that since the computers are down UAL can't book pax onto any other airline ..."
"Systemwide Ground Stop posted at FAA: Due to USER REQUEST DUE TO AUTOMATION ISSUES. UAL AND SUBS ONLY., departure traffic destined to ALL airport will not be allowed to depart until at or after 13:15 UTC."
I was guessing something a bit more
01 OUTAGE-DATE-DATA.
05 OUTAGE-DATE.
10 OUTAGE-YEAR PIC 9(04).
10 OUTAGE-MONTH PIC 9(02).
10 OUTAGE-DAY PIC 9(02).
05 OUTAGE-TIME.
10 OUTAGE-HOURS PIC 9(02).
10 OUTAGE-MINUTE PIC 9(02).
10 OUTAGE-SECOND PIC 9(02).
10 OUTAGE-MILLISECONDS PIC 9(02).
05 OUTAGE-DIFF-FROM-GMT PIC S9(04).
MOVE FUNCTION CURRENT-DATE TO OUTAGE-DATE-DATAI always kinda liked the meme from that language of "variable as picture". Calling something a picture has inherent and obvious implications WRT call by value vs call by name.
On the other hand intrinsic functions like FUNCTION CURRENT-DATE are not as logically amusing.
The United States has the world's largest economy, according to the IMF, World Bank and UN. So obviously the US is doing something right. Perhaps the rest of the world ought to adopt American conventions. I say that in jest, but the point is that criticizing the American way of doing things is a popular sport, yet at the end of the day, the American system has resulted in an economic output greater than Germany, the UK, France and Italy combined. The EU has a per-capita GDP of $36,779 and the US is at $54,601. So apparently something is working with the American way of doing things. Just because "the rest of the world" does it doesn't make it better. This whole idea that the American way of doing things is somehow inferior is just as ridiculous as the British/French rivalry.
The fact that the parent comment is the top comment is ridiculous. It's a petty thing about which to complain and adds nothing to the discussion. In fact, the parent comment ought to be down voted for being absolutely irrelevant to the posting in question.
https://en.wikipedia.org/wiki/Ground_proximity_warning_syste...
https://en.wikipedia.org/wiki/Traffic_Collision_Avoidance_Sy...
https://en.wikipedia.org/wiki/%C3%9Cberlingen_mid-air_collis...
https://en.wikipedia.org/wiki/Tupolev_Tu-154#Incidents_and_a...
I expect there's growth on all of these:
- Total planes
- Total flights
- Number of routes
- Number of passengers
- Passenger utilization of electronic boarding
- Passenger utilization of reservation modification (seat changes, upgrades, etc.)
- Passenger utilization of in-flight electronic amenities (wifi, in flight entertainment, etc.) that are billed through the system (United has a proprietary WiFi that's tied to your frequent flier account)
- Online booking
- Travel agent/reseller pricing queries
- Travel agent/reseller booking
I'm not saying any of these are dramatically increasing, but I'd bet they're all going up slowly, which will add more and more load to the system.
source: http://www.businesstravelnews.com/Travel-Management/Drop-In-...
Yes, one is wise to attribute to incompetence over malice, but these systems are demonstrably not run by incompetents: they've been in operation for decades.