The real issue is that a single erroneous flight plan should not stop the full system.
The real issue is that a single erroneous flight plan should not stop the full system.
First, you don't understand the purpose of a flight plan. The purpose of the flightplan safety.
You, the captain, are saying "My intention going from A to Z via K,J and X". You then "activate" your flightplan over the radio upon departure.
The purpose of the flightplan is so people on the ground broadly have an idea of where you are and when you'll arrive. The former is to assist with co-ordination between control zones. The latter is so that if the destination airport doesn't see you X hours after your ETA, then they can start searching along the filed route until they find the debris of your crashed aircraft.
Second, you need to take "automation" out of your thinking. Flight crew don't just punch the plan into the nav and sit there twiddling their thumbs.
Shit happens and the crew might need to deviate from the plan. It happens all day every day. There might be some horrible weather ahead that you want to route around. There might be something happening in the airspace around you that ground might want to route you around.
Therefore flight plan merely describes your INTENTION.
To make the system resilient.
Note: When the error happened the flight was not in the air yet. The system in question received the flight plan 4 hours before departure time. If the system would have flagged the flight plan as bad they could have called them and told them that they can't fly. If not that they could have refused them when they were entering the airspace. Can happen any time for any reason.
> Naïvely it sounds like having even a single plane in an unknown position would make safe automated control impossible.
Flight plans are not for knowing the position of the plane.
> Is that wrong?
Slightly.
It sounds to me like the engineers made a design decision between "add handling mechanisms for valid-but-unexpected flight plans" and "ensure we can handle absolutely every valid flight plan"? If so, this is a rare case where I sympathise with the engineering team behind a major IT failure.
Flight crew and dispatchers report that flight plans are regularly rejected, and then they need to file a new one. But it sounds like this rejection happens in a layer before the one which failed in this case.
So this is basically a system which is not meant to be able to reject a flight plan, since the plans it receives were already checked and validated.
The issue should have been caught in one of the higher systems, and then the error should have also been handled in a more appropriate way.
This is not true. The plan is transferred 4 hours before it is due to enter UK airspace. Big difference. Flights can already be in the air when the plan is transferred.
You are correct.
Remember anyone flying IFR needs a flight plan, lots of private pilots etc. Plenty of people make mistakes.
(Hint again: Something certainly could have been done much better there in this system, but the limited component itself that all caused this maybe even just had no other option at that point. Been there once in those many component systems talking with each other across with too many interfaces, people, etc involved, part of them safety critical and also subject to endless requirements and certifications? Can recommend..)
But the flight plan wasn't erroneous - it was perfectly valid. It just happened to have a weird coincidence of two identically tagged airports spanning the UK. If someone had put an extra waypoint in between the two, nothing would have broken (AIUI).
Well. Then the real issue is that a single flight plan, erroneous or not, should not stop the full system. It should flag that flight plan for problem resolution and keep on working with the rest of the flight plans.
Would result in a slightly inconvenienced flight dispatcher, or worst case an inconvenienced flight, instead of a nationwide shutdown.
Sure but "flight data is safety critical information that is passed to ATCOs the system must be sure it is correct and could not do so in this case" is a reasonably valid stance in a super-critical system that must absolutely defer to safe operations in the case of error.
That is not reasonable. And it won't become reasonable by alluding to how "super-critical" or "safety-critical" the system is.
Stopping in the face of being unable to verify that safety critical information is correct is a perfectly good choice, especially when the good consequences are fairly light like "some people are a bit late with their flights" and the bad consequences may be 10s or 100s of people dying.
I think this might be the root of the misunderstanding between us. The consequences of an airplane not having a flight plan is not “10s or 100s of people dying”. The consequences of rejecting a flight plan are:
- if caught on the ground:the one flight is delayed. They don’t take off until they manage to file a fligtplan the system accepts.
- if caught at the airspace boundary when they arrive without a flight plan in the system there are two options: if ATC want to be hard-ass they will be refused entry and have to land at an alternate airport. If ATC is acomodating they will be asked at every new controller “<flightnumber> state your intentions”.
Either way the system is fine and dandy. This is not what keeps airplanes from coliding with each other or the ground.
Obviously the first option, catching the problem on the ground, is better because it is not increasing the workload of the ATC personel.
Here is the thing, you are talking about “the system cannot verify this flight plan as being correct". And you are right! It cannot. So it shouldn’t try. It should give up. On that flight plan. I totally agree with that. But there were lots of other flight plans it could and should have verified.
Have you done web development? What would you think of a server where a single weird edge case request would stop the whole service for hours? Would you say it is well engineered?
The consequence of not being able to route a flight through reasonably congested airspace safely may well be "10s or 100s of people dying".
> when they arrive without a flight plan in the system
They had a flight plan but the system could not verify that it had correct data. i.e. the system believe it had internal corruption and should shut down.
> What would you think of a server where a single weird edge case request would stop the whole service for hours?
What would you think of a service where internal checks suggested corruption and it kept on serving potential corrupt data to customers and writing it to databases? I would consider that a failure in extremis.
And I would agree. So stop serving the potentially corrupted data to costumers. Mark it as potentially corrupt and let an operator handle that potentially corrupt data manually while the system keeps processing the the rest of the not corrupt data.
Is this such a hard concept to understand? Because it feels like it is not getting through.
I'm sure no doubt they will now fix the issue, but you can only aim to handle faults that you predict during development, plan for unexpected faults, and hope that something super unexpected doesn't cause cascading faults
Are you sure? Have you also consulted with zimpenfish? From their comment it seems they disagree.
> hope that something super unexpected doesn't cause cascading faults
But that is the thing. Processing failing isn’t and shouldn’t be “unexpected”. Here the individual processing items are flight plans, but the concept is more general. They could be “requests” to a web server, they could be “tasks” in an async worker node, they could be “transactions” in a database, if you do any compute with a work item it can fail. You need to put a big “try-catch” around it and add a field in your database or queue or remote procedure call system to mark them as failed.
One work item failing shouldn’t be “super unexpected”. Even if you can’t name the exact conditions under witch it might fail.
They might well have considered it the possibility of duplicate tags in the same flight plan but since identical tags are supposed to geographically distinct, they probably didn't further consider a flight plan where the entry and exit points from the UK (which is a relatively small airspace, I think) are the same tag (since it sounds from the report they did that this was an extremely weird flight plan.)
I don't see it said anywhere that the two same-named waypoints were consecutive, with no other points between them. Just that they were "approximately 4000 nautical miles apart". Or that being consecutive duplicates rather than just duplicates was the cause.
Since they don't name the waypoints, it is very vague.
[0] https://publicapps.caa.co.uk/docs/33/NERL%20Major%20Incident...
No waypoints that enter UK airspace between them, since "In this specific event, both of the waypoints were located outside of the UK". (1)
So there could be any number from zero up, of international waters waypoints between them?
1) https://publicapps.caa.co.uk/docs/33/NERL%20Major%20Incident...
For me, it read like "there was no route between Aa and Ab because there were no waypoints available between them". But yeah, without a more detailed breakdown, information, and/or a better understanding of the precise format of the files and procedures, it is going to be a bit of assumption and guesswork.
> two separate waypoints which – although geographically distant – had identical designators.
Also to be honest, shouldn't this have been better unit tested? Considering aviation is a ~ £60 billion industry for the UK.
Entirely possible there was a unit test that confirmed the system would error out in this particular condition. This sounds more like a requirements issue.
Relying on it's unit tested strong systems does not make. There's a lot more to testing.
We don't know how well it was tested. (unit or otherwise)
What we know that it has processed 15 millions of flight plans previously without revealing this flaw. So they must have got some things right.
Sure, the unexpected identical naming should not cause an issue, but when you're developing software like this you conduct multiple risk analyses and FMEAs to attempt to design a system that can handle faults that you can predict
You can have failovers, safe states, recovery procedures up to the eyeballs, but if a situation occurs that you had not predicted then you're in unknown territory. Your system may handle the fault and fail safely, or it may cause an unexpected chain of events.
TLDR: You can plan for faults all you want, but you can never guarantee it's 100% bombproof