Whiteboards used after Gatwick flight information screens fail
bbc.com
bbc.com
The US airline's apps tend to be very good, sending you real time notifications of gate changes and flight delays. On Delta, the app even took me directly to a rebooking function when it notified me of a flight delay during a layover in ATL. Very good UX.
Jetblue is incredibly sensitive about any change and will do a push notification for it. (Your flight landed 2 minutes late, etc)
Other times? Yeah...not so much
(they're better when you enable push notifications for your flights (and I set up SMS notifications, too))
A big improvement has been wifi availability and stability on the plane. Not so much because I can browse the web in flight, but because regardless of whether I get in flight wifi, the airline's app uses it for free. So now that wifi is reliably always there, I know what's up and where I'm going before I walk off the plane.
I have found, however, if you sign up for SMS updates through most airlines they appear to be quite reliable, and sometimes beat the in-airport departure boards by several minutes, particularly for gate changes.
Oddly enough, Gatwick's CIO was last week [1] telling journalists that he didn't like to rely on the cloud:
"The airport has used multi-tenanted public cloud in the past, but for that to work perfectly you need your telecom providers to work well, the path across internet to work, the hosting firm's telecoms to work, and their data centre support and monitoring not to drop. We have had examples where even in the public cloud using dedicated slices, you can suffer from noisy neighbours, and we just can't have that. Our operations are too important to us, so we need to keep it close. We use cloud for resilience, and we're very pro-cloud, but not for core services."
So it seems odd that departure boards were being fed from a cloud service. Maybe they hadn't been considered as a core service until now.
On the bright side at least the failure was limited to just the boards. The website, apps and check in desks all had the correct information.
[1] https://www.computing.co.uk/ctg/news/3034002/too-many-compan...
Techs would mark down "NG" for no go, OK for ok, and HOLD...these all make sense to us, but not to customers. Even worse, once an hour or two a tired tech would shuffle into the waiting room study the board for a minute, and slowly cross out the name of a person. The whiteboard worked, but that day our shop was indistinguishable from a mafia den.
If the airport has a dedicated fiber channel between their data center and airport itself and it was damaged, this could be the result. The question I'd ask is why they have no fallback strategies (e.g. cellular, VPN over a less secure internet channel, etc).
This might be an area where they can explore graceful degradation, meaning that gate information would continue to display (from airport ground control), but some flight times might be missing.
I fully agree with you on your last sentence. One should wait for a bit more than a vague declaration reported in a newspaper before giving unsolicited advice.
How many people (even tech ops profesionals), these days, know how to tell the difference? Of those that do, how many don't just keep their mouth shut because their bosses don't want to hear it?
These details are just trusted to "the cloud", instead.
Of course, the airport itself isn't a datacenter of any tier, and that "last mile" can be the biggest redundancy challenge, anyway.
The medium is the message, etc...
If there was a digital fallback, it would absolutely have to have something to tell commuters that the information might be stale (as someone frantically types up flight updates in a back room somewhere)
They do. It's the whiteboards. Whiteboards are cheap, resilient and robust. Exactly what one wants for a backup.
That said it was probably very expensive :P
Theres a lot of fawning over the humble whiteboard here, just wanted to emphasize that they were part of a larger scheme
But why does the data link itself have no fallback? An important data link should have multiple dissimilar backups.
It's Gatwick, not Heathrow. Whiteboards are good enough. If every airport the size of Gatwick (and busier) implemented digital fallbacks, I wouldn't be surprised if the fallbacks took out the primary system more often than external causes. (I can’t remember the last time I used those screens over my phone. In a certain sense, the screens are now a backup.)
More precisely, if you consider the number of situations which (a) take down the primary system but (b) don't take out the airport, and try to cover those as cheaply and robustly as possible, for all but the world's busiest airports you come up with whiteboards. Localised power failure? Whiteboards. Surge? Whiteboards. Limited fire or flooding? Whiteboards. Fibre-optic cable gets nicked? Whiteboards or alternative digital system. Now compare the costs and risks of the digital system. Whiteboards are the work of a smart operator who probably saved their airport years of technical debt.
The evidence is that whiteboards are ok, but with their location a second fibre link would not be difficult.
“High availability” database measures are often less reliable than the normal, simpler approach.
Automated network failover equipment often causes more downtime than having multiple links and only failing over manually, or even depending on a single link.
Complexity is the enemy of robustness.
System failure is not just one simple category with known behaviours; different types of failures affect performance in different ways, and it’s hard to support all of them well. Just as exception-handling code is commonly the buggiest code out there, high-availability measures are commonly some of the most fragile parts out there, unintentionally causing more problems than they solve. They can be made robust, but it takes a lot of effort to do so, and frankly it’s seldom justifiable. Typically, simplicity is good. Definitely it’s easier to debug.
Failing over manually sounds fine. If they can't swap over at all then that's a sign of fragility which will cause its own problems.
https://www.computerweekly.com/news/252441044/Gatwick-Airpor...
[Gatwick has launched an app] for internal staff at the airport, such as shop assistants, baggage handlers and train station employees. ... He says the need for the app stems from some passengers entering “panic mode” and expecting to get flight information from anyone who appears to work at the airport. They may approach an employee in a non-passenger-facing department who has no more information than they do. ... “Passengers sometimes get more information from Twitter and Facebook feeds than from staff on the ground,” says Chacko.
“This is an attempt to bridge that communication gap and make everyone look smarter and more up to date.”
I wonder how well it worked today
https://www.computerweekly.com/news/4500269843/Gatwicks-Airp...
Aesthetically it was so much better than big super bright screens.
Sadly, they had to go digital in the end due to spare parts becoming difficult to source.
According to Wikipedia, there are still some airports (eg Frankfurt) and many rail stations that still have them in active service.
Not many companies have 10 spare people on site ready to act as emergency whiteboardmen.
They'll stare at the blank screens and wonder what's happening. A whiteboard in front of those screens, and other hot-spots, is far more considerate of how people behave inside a departure lounge.
Low-tech solutions to high-tech problems can be quite beautiful when executed well. It's a sign of good contingency planning because they anticipated no network availability, rather than betting on network redundancy.
I hope the few people who missed their flights were compensated in some way (refunds, seats on the next available flight...).
Checkin generally opens 24 hours prior to departure (some airlines now allow >7 days prior) - hence for the vast majority of passengers the gate is unavailable as ticket issuance time.
Most airports and airlines let you get gate information fro their websites/apps as well though..
It seems like a lot of work people do is individual and idle when other systems take responsibility of many issues in a process. I'm glad that people worked closer together as a result of this failure!
I'm glad people worked closer together here, not that there was a failure!
https://www.telegraph.co.uk/news/2017/05/13/cyber-attack-hit...
Wonder why someone didn't just prop up PC's with large monitors or projectors around the airport, hooked into the WiFi to broadcast this info?
It's simply that the backup plan to the failing digital boards, are plain whiteboards.
The mobile apps are great when they work, but still need to be considered a convenience rather than a primary information source.
Corollary:
A few do have a phone, but it is not a smartphone.
In my experience, changes often show up in Kayak ahead of the airport dashboards. I check apps or notifications before the official dashboards now.