A400M Airbus Flier crashed because of software issues
translate.google.com
translate.google.com
All safety critical software (every piece of code ran on-board is safety critical the least) in aerospace needs to pass the DO-178 standard [1].
That is far more serious than standard unit tests you are used to in node.js applications. Generally speaking, to develop a piece of code under that standard it takes 20% of time to write the code, and 80% to testing, and enormous amount of documentation (that is optimistic estimation, usually worse).
Quoting speaker from DO-178 training course I attended:
People often ask us. "How do we know the standard works?"
We give this answer: "We do not know. But there have been zero crashes due to software issues since introduction"
If this crash confirms the cause to be a software bug, that is something much bigger than an airplane crash - a huge punch to the whole federal aviation administration.However it's a complete different matter for civil airplanes : it's the lead dev/project manager which is liable for life for what he has shipped. For example, a retired engineer from Airbus was heard in trial for the Concorde accident in France in 2000.
I wonder if someone can pull an #ElonMusk on the DO-178 to slim things down in order to have better control.
ps: planes fly with bugs, see the DreamLiner, Airbus ones aren't free from them either, employees know this. (I guess they travel by train)
The train brake control systems are written in ... C.
For a quick example, there are many languages where you cannot accidentally run off the end/start of an array, barring a compiler error.
With a language like C / C++, it's possible. Not probable, given that sort of testing. But possible nonetheless.
Some languages are also easier to test than others, partially because of this, partially because of other issues. For instance, in some languages you can guarantee at the language level (again, barring compiler errors) that something won't be modified. (Like const, but actually working.)
On the other hand, I also wonder why people don't think of using Erlang for something like this. The VM is designed for a ridiculous amount of uptime, it has a supervisor tree that can monitor and restart failed processes, and it can interface with C/C++. A rock-solid VM should be running and monitoring life-safety systems.
What alternative are you suggesting?
AFAIK (partial) assurance in C/C++ can only be handled by additional testing tools, Frama-C for instance.
I agree that C/C++ should not be used for security applications. Ada is a much better choice because it was designed for security.
In fact, they (http://windriver.com/products/product-overviews/PO_Diab_Comp...) say:
Diab Compiler has been a reliable code generation tool for
avionics products certified for DO-178B, products for the
nuclear market certified to IEC 60880, railway applications
certified to EN 50128, and industrial products certified
to IEC 61508, and is now qualified for use in automotive
applications certified to ISO 26262.
Ada does have some built-in advantages, but I think my point still stands: the language is a small part of the entire SDLC, and I don't think it's the most important part.I'm pretty sure that the only industrial formally verified compiler is CompCert (for C), though I could be wrong. The motivation for CompCert was certainly that Airbus wanted such a compiler.
Ada wouldn't be a better choice simply because it's designed for security. It'd be a better choice if it turned out better in practice. I've read some of the studies that have been done, and I haven't found them convincing.
Requiring additional tools just isn't a problem, if it works well. Don't criticise the process, criticise the result.
The focus on languages instead of the SDLC is telling, I think.
I'll speak to the 787 avionics system. It used a system of channels/buffers and processes very much like what Erlang and Go use for interprocess communication (I'm trying to remember now if channels could be received in multiple processes like Go or if only a single process could receive like in Erlang). This was an excellent model for what we were doing, and really for a lot of systems this sort of CSP and actor style model maps well.
I once worked on satellite systems using Ada95 and Ada2005 (Ada is definitely not dead). The language is a pain to use but is impressive in that it catches more crap at compile-time than anything else I've seen.
True. More information: http://www.ada2012.org
The actual failure rate of software to this standard seems to be at least two orders of magnitude higher.
Formal verification and state space analysis can prove that the software "model" will not fail. State space exploration of the actual implementation is actually often not feasible due to the enourmous amount of states.
So my question: Are you doing formal analysis of the software models/designs? I know they are using it in software used in space crafts / nuclear power plants.
Reference: - http://ti.arc.nasa.gov/m/profile/dimitra/euromicro-share.pdf - http://javapathfinder.sourceforge.net/
For example Astree [1] has been developed for decades now, with Airbus as one of the major sponsors.
The standard does mention "formal verification/methods/proof" but to my knowledge it's rarely been used extensively.
What made it better:
Concurrent editing. Only one person at a time was modifying any section, but two or more sections could be edited by a group of people.
Version control baked in. VC is awesome, reducing friction means that it actually gets used. Checking in/out sections was just part of the process.
Automatic traceability matrices. These are not easy to make by hand. That way is error prone and tedious. It's also hard to verify. With DOORS we were able to generate these automatically. Verifying them was relatively painless because we weren't flipping through several 1k page documents to make sure things actually matched, we could easily skip to only the applicable parts of any document. The main error that it didn't reduce was when a requirement didn't get linked to every test case or design feature that it should've been linked to.
Again, probably better tools out there than DOORS (at the time, or hopefully by now). But it's a lot better than what most offices do with either post hoc document generation or ad hoc generation with Word + Excel.
Sure sounds like the way it should be to me. Maybe it needs to be made easier to prove your software is correct, but to me it seems like for systems like airplanes, code that cannot be proven correct is worthless.
Considering most programmers are said to produce 6 lines of good code a day, maybe it's not even actually slower in the end if the formal verification process filters out every other line you would have written that day.
the apologetic attitude needs to stop. programs are just complicated machines and should be treated as such not some mystical voodo that will be broken no matter what
Airbus Defence and Space has today (Tuesday 19 May) sent an Alert Operator Transmission (AOT) to all operators of the A400M informing them about specific checks to be performed on the fleet.
To avoid potential risks in any future flights, Airbus Defence and Space has informed the operators about necessary actions to take. In addition, these results have immediately been shared with the official investigation team.
The AOT requires Operators to perform one-time specific checks of the Electronic Control Units (ECU) on each of the aircraft's engines before next flight and introduces additional detailed checks to be carried out in the event of any subsequent engine or ECU replacement.
This AOT results from Airbus Defence and Space's internal analysis and is issued as part of the Continued Airworthiness activities, independently from the on-going Official investigation.
They're asking for a one-time check to be performed on the ECU. If it's just a software bug, normally a one-time check wouldn't reveal whether or not that bug could trigger. So, obviously they've found something they're concerned about, but it seems to me to be a bit early to say as Spiegel Online do that software caused the crash.
In any event, one of the flight recorders was only just sent off to the manufacturer: http://economictimes.indiatimes.com/news/international/world...
As far as JIT manufacturing, that's certainly the case for bigger parts and systems, but it sounds like the ECTs are replaceable so they probably have backup replacements stocked near most major airports. (Assuming it's a part that can be replaced during routine pre-flight maintenance.) And the components that go into the ECTs are almost certainly produced in large batches rather than continuously.
I would assume the Spiegel report is based on more sources than just Airbus' public statement. At least most of the time, Der Spiegel is a real newspaper, not just some blog spouting stuff based on conjecture from a single source. They mention internal sources in multiple places in the article.
I will wait for the official accident investigation report.
The media is infamous for not getting facts straight. They much rather write an opinion based on suspicion.
Well, no, thank you. Give me facts, get them straight, then I will make up my opinion.
We have had a long time to recognize that our brains don't work well. It is time to accept the facts.
There is a critical time period during a take off when the aircraft is at maximum risk. If an engine fails before rotation (i.e. before the wheels leave the ground) an alert crew can stand on the brakes and use thrust reversers. The aircraft may get dinged up, but there is a reasonable chance the crew (and passengers) will survive.
If there is an engine failure after rotation but before the aircraft has gained sufficient altitude, unless there's a big, flat field next to the runway a crash is almost inevitable.
When an aircraft turns, it will lose altitude unless the crew compensates by adding power. An aircraft without power and sufficient altitude cannot make the turns necessary to go all the way around to land on the runway they just left.
There are three relevant speeds for large aircraft. (I'm going to generalize very slightly to keep this short and readable.)
V1, Vr, V2.
V1 is the takeoff decision speed. An engine failure recognized before reaching V1 is handled by aborting the takeoff. An engine failure recognized after reaching V1 is handled by continuing the takeoff. At the V1 callout, the pilot flying removes their hands from the top of the throttles (as a physical reminder that aborting/rejecting the takeoff(RTO) is not happening for a simple engine failure).
Vr is the rotation speed, where the nose wheel is lifted from the ground.
V2 is the speed at which the airplane will climb safely with one engine INOP.
In most cases, V1 is the lowest speed, meaning there are cases (between V1 and Vr) where an engine out with the nosewheel on the ground results in continued acceleration, then rotation, and flight.
It's a checkride bust to RTO above V1 for a simple engine failure.
I did not know that there was a situation when a flight crew would continue a take off after an engine failure but before rotation.
Why should it be?
The attrition rate is terrible and its impossible to keep people who have knowledge of the program on board.
I would guess that 30-50% of our team are inexperienced java software engineers with less then 5 years of working experience. Around 100 people work on our project but there are probably 3-5 people who still have a clear picture of how everything is working / supposed to work.
and if the contractors are from India, the competition does not stop after the contract have been awarded. During the execution of the project, the competitive behavior will continue. One contracting firm will compete with an aim to reduced the tasks/responsibility of the other contracting firm, thus (or hoping to) capturing more contracting hours. Very vicious and dangerous..
I still don't understand "The crash is the worst accident since the development of the A400M", though. To me that implies there was a pretty devastating accident during the development.
I think it's sloppy journalist writing for the worst accident that happened with the development of the A400M.
Two translated quotes: "The investigation yield a clear result: Shortly after the lift-off of the test machine, the computers send conflicting commands to the three engines which then powered off."
"Soon after the crash, experts of the German Air Force suspected a software issue with the fuel supply unit because such a fatal drop of power so soon after the start could hardly be explained differently."
So, not much information why the computers sent conflicting commands and also why the engines power down in such a situation.
I think shutting down the engines is probably the safest option when this sort of thing happens. You could argue they should stay in the present setting, but what would happen if one engine were at 0% and another 100%?
Most aircraft are pretty good at gliding even without power, and I'd assume a deadstick landing is part of the pilots training. In 2001, TS236 flew unpowered for 19 minutes before making an emergency landing (on a runway) with only minor injuries:
Part of earning your multi-engine cert is memorizing all manner of different airspeed limits for that kind of situation. If you want to maintain yaw control with one engine feathered (er, shutdown) and the other at full throttle you must be going faster than X knots or whatever. Below that indicated airspeed you pull back on the throttle or you're going into a turn at best or more likely a spin.
Planes are designed to fly fine in that situation - it's what you get if one engine breaks down. Now landing the thing with one engine stuck on 100% would be interesting. I guess you could kill the engine somehow - turn off the fuel or pull the fuses.
https://en.wikipedia.org/wiki/Qantas_Flight_32
Mind you, that was the result of an uncontained engine failure.
Shutting down an engine should always be a decision made by the pilot not a machine IMO. A pilot might prefer to blow out an engine if that gives him enough ( even a couple of seconds matter in this situation ) time to reach a save landing spot or avoid an obstacle.
> what would happen if one engine were at 0% and another 100%?
Pilots are train in this situation all the time and is part of the syllabus for a multi-engine rating. Basically the plane would try to turn to the side that the engine failed. Pilot will use opposite ruder and aileron to compensate while cutting back power in the good engine to just enough you can keep altitude if possible.
> Most aircraft are pretty good at gliding even without power, and I'd assume a deadstick landing is part of the pilots training.
So happen here too. The crew tried tried a dead-stick landing in a field when they realized the could not make the airport. Unfortunately the hit a High Power pole and the plane catch fire.
One of the crew members lost in the accident was a friend of my father. My condolence to the families of those who lost their live that day.
What is there to lose by opening up the software to criticism other than better aviation safety? We know that obfuscating / hiding source code does not make applications / platforms safer or less at risk to malicious behaviour so I'd like to challenge the manufacturers to do so.
The other problem is that the maintainers have to be set up to handle a potential avalanche of comments, criticisms, questions, and pull requests, mostly from people who don't know anything about software development processes and standards within the aviation community. If they're already too overloaded to find all of the bugs themselves, they certainly won't be able to effectively manage open-sourcing their code.
I mean, who would fix problems if it's just a hobby? And if it's open, it must be a hobby. Surely, that can't possibly work!
Open sourcing something like Linux works very well _precisely_ because it has a very large audience and is (relatively) approachable by hobbyists too.
On the other hand, aerospace engineering and software is narrowly specialized with a (relatively) small group of experts and code used in commercial/military aircraft is anything but approachable to hobbyists.
Then there is the fact that unit testing this kind of code requires engineering knowledge of the specific hardware involved (e.g. not just any jet engine, but one very specific model). Finally, let us not even mention the huge pink elephant in the room, namely that the absolute and vast majority of "hackers" does not have access to jet engines used in commercial (or military) airplanes and even fewer have the ability to conduct test flights.
But this is different. I wonder if they need to rethink their approach to software? Four engines, running the same software --> single point of failure.
"Problems in developing the engines, and particularly in certifying the engine control software, contributed to three years of delays and a new cash injection by governments in 2010."
Seems like they had some issues. Surprising it's that difficult.
http://www.reuters.com/article/2015/05/19/airbus-a400m-idUSL...
http://www.bloomberg.com/news/articles/2015-05-18/hacker-cla...
Unfortunately manufacturing let a batch of 0.0201 size slip thru inspection, they just barely fit on the assembly line despite being out of spec too large, and the PID controller makes the system go into oscillation and explode because its outside its theoretical limits of whats possible.
The most insidious spec violations are "manufactured too well" if for example you rely on frictional damping to eliminate oscillations, then making and shipping better ball bearings than you'd ever shipped before, could ruin an overall system because one component is too good. Possibly the "error" is something is too smooth, too straight, too flat, or too low friction. No one ever expects those to cause a disaster, but it can happen.
That would be an example of a disaster involving software, that can be fixed in software, although it wasn't caused by the software, it was caused by bad control system engineering design work. Also this abstract example probably has nothing to do with the real problem, although the "widget" and "size" could very well be something line fuel line tubing inside diameter.