Frankfurt rail work damaged fibre, causes global Lufthansa check-in system fault
dw.com
dw.com
*Please*
All critical system shall be working autonomous. Cache flight/passenger data locally at airport for next weeks. Servers are for synchronizing, not for keeping your data away. With a local cache you can keep going for some weeks, adding/changing bookings will be harder but you can remain in the air. That is pretty complicated and requires more work but is important. Git or IMAP are examples how it should be.
Outages will happen and will keep happening and when the get more seldom, the will become more serious because it will hit inexperienced staff. Especially British Airways is known for issues[2][3].
[0] https://www.lufthansa.com/xx/en/flight-information.html
[1] https://www.heise.de/news/Sabotage-bei-der-Bahn-Viele-vertra...
[2] https://thepointsguy.co.uk/2017/09/amadeus-network-issue-cau...
[3] https://www.datacenterdynamics.com/en/news/british-airways-e...
PS: Deutsche Lufthansa is known for reliable transport. Deutsche Bahn is known for unreliable transport.
I think a more workable solution would have been to treat transit for this data center how most other data centers are organized, with multiple independent uplinks that take different routes into the building. A last-resort point-to-point wireless link would probably be wise as well.
That is enough for some hours and even days. If you're going advanced you can allow to add (right at the gate/counter) one person to more passengers and so on. You can also go all-in and add code which allows merging of external data i.e. allowing headquarters and airport personnel working at same time.
Die Fluggäste wurden auch nicht per Strichliste in die bereitstehenden Maschinen gelassen, weil nach Angaben des Personals wichtige Informationen zum Abflug fehlten.
If it works that is an appropriate solution.PS: We cannot see circumstances. Maybe a hack, a burning datacenter, war, co-worker going crazy, international sanctions, storm, flooding, whatever. For example Deutsche Bahn itself suffered recently a sabotage, two entirely independent glass fibers get cut off at same time (one in west-germany, one in east-germany).
German source: https://www.faz.net/agenturmeldungen/dpa/it-probleme-bei-luf...
They cut off four cables at once. It didn't failed immediately.
An excavator. One.
That's not diversity. Diversity would require cables cuts at two separate sites.
[0]: https://twitter.com/deutschetelekom/status/16258656661117091...
Technically it wasn't the Deutsche Bahn, but a hired construction company. And they just happen to cut some cable from Deutsche Telekom, which also broke the connection from Lufthansa and some others. Those things happen with constructions. But generally, they all are kinda at fault.
> All critical system shall be working autonomous.
It's not always obvious what is a critical system. I mean it was a semi-important sub-system of a bigger company, not something where you would assume it could cause a national impact.
Why would you not assume it was something that could cause national impact when cutting through fiber cables was the very thing that impacted rail travel for Deutsche Bahn 4 months ago in Northern Germany?[1]
DB uses GSM-R(Global System for Mobile Communications Railway)which has fiber backhaul at regular intervals along the tracks.[2]
Given the incident in October is still fresh in memory surely there should have been a heightened awareness to this.
[2] https://www.globalrailwayreview.com/article/1088/current-gsm...
Article says that Lufthansa actually has a backup lane and it initially worked. It finally failed later [sic!] on Wednesday, maybe because the load was too much for the backup lane. So they tried - as usual - to lower the chance of an outage. But it is crucial to actually being able to keep working with an outage.
PS: The construction workers seem to have drilled through four cables (each ~ 900 fibres) in 4-5 meter deep below the surface.
No, Lufthansa not having sufficient resilience was to blame.
Hopefully they fix it until next Monday. I'm having a flight to Frankfurt then.
Off topic, sorry. I'm curious why Germans often use this phrasing. Is there a "false friend" [1] in English? (I assume you mean 'before', not 'until')
Don't go there before Monday. Don't go there until Monday.
You can get this deal until Monday. You can get this deal before Monday. (Some lack of clarity about whether the offer is valid on Monday with until though)
But I agree, it doesn't work in this context.
English is weird.
https://dictionary.cambridge.org/grammar/british-grammar/adj...
It is incredibly common for multiple fiber bundles to have a shared path, because it's often easier to locate another bundle in the same place. Fiber is often laid along rail lines because rail lines have the perfect property shape for communication, and there's specialty train cars that make laying fiber alongside the rail really efficient.
No, it's not. Metro Fiber links are designed and deployed in a ring topology for exactly this reason. See:
It's really easy for a vendor to claim that there's no point at which a single backhoe can take out two of the A/B/C/D/E paths in your diagram. It's also pretty easy to buy dark fiber from two different vendors that are reselling in the same bundle or have their bundles placed in the same conduit.
>"It's also pretty easy to buy dark fiber from two different vendors that are reselling in the same bundle or have their bundles placed in the same conduit."
Wrong again. If you are purchasing an IRU on dark fiber you know exactly who owns the physical assets. If you are purchasing wavelengths from a reseller they will happily disclose whose network they are reselling. Additionally you can request the CLR/DLR for your circuit and see exactly how it's built. You can also easily avoid resellers and not worry at all about this.
Lastly you seem to not understand the difference between long haul and metro fiber.
edit: related tweet from Deutsche Telekom
https://twitter.com/deutschetelekom/status/16258248409249505...
It is good to prevent failures. But autonomous local systems shall remain (basically) usable for some time. And if the local system is a paper, that's at least something. We've also ABS, ESP and ASR in cars. We still close the seatbelts.
This is 100% on poor management practices by Lufthansa, but of course they’re going to point the press to that shiny object over there in the form of a fiber cut. The press, as usual, took the bait.
Had there been multiple cuts in multiple different locations I'd be more sympathetic. Shetland being a recent example [0], where one fibre was cut, then the other one in a completely different direction was also cut.
Assuming they had 2 cables and both are broken, whether something as critical as the entire LH booking system deserves more resilience than just 2 geographically diverse cables, I'm not sure -- ultimately it's what's the cost, what's the damage, and what's the likelihood of it happening.
[0] https://www.bbc.co.uk/news/uk-scotland-north-east-orkney-she...
We switched from microwave antenna which has its own issues to fiber and my dad is thinking, "Well, you said this would be better." You can blame the local ISP, but am I wrong to think that it's hard for a rural ISP providing fiber to afford redundancy?
Though, it's also possible that the other connections just failed, or could not take the sudden traffic, or the system crashed for some nonsical reason, because nobody really tested this scenario.