Patriot missile software failure, 28 soldiers died. Fix: reboot the system
en.wikipedia.org
en.wikipedia.org
The "software error" Wiki alludes to is that the Patriot missile kept track of its internal clock with floating point numbers. When the machine had been booted in the recent past, such as every time in testing, the floating point number spent most of its precision to the right of the decimal point. This let it able to do the designed behavior, which was calculate very small delta(time) to be able to do velocity/position calculations and get fairly close to fast moving objects then go boom.
The problem is that floating point numbers have a limited amount of precision available to them, and if you are using a few billion milliseconds (2 weeks), almost all of your precision is lost to the left of the decimal point (and, given that this is precision-intensive work, you didn't need to wait that long to see anomalies).
Lower precision meant that taking delta(time) got increasingly less precise as time went on. Which meant that velocity/position calculations got progressively more screwed up. Which meant the missile did not go boom in the general vicinity of incoming missiles. Which killed Americans and allies.
Thus the moral of the lecture: a) your computer is a powerful, tricksy beast which has many ways to trap you in even straightforward code and b) you should treat software quality like some 19 year old's life depends on it, because it might.
You should treat it like that if someone's life does depend on it and you have the resources to develop accordingly.
If you're developing something like, say, bingo software, you're probably better off devoting time to improving the product or marketing it, rather than working on it being 100% bug free.
It's all a tradeoff - time spent on eliminating every last little bug is time not spent on adding features or making it faster or marketing it or whatever.
Edit somewhat less clear-cut cases might be bits of software that you release publicly, and subsequently get used for life-critical systems. However, in that case, the onus is on those adapting the code for use in that environment to provide the testing/review/etc... rather than blaming the upstream developer.
That doesn't mean it's your fault if someone uses your free XML parser in an amusement park ride and your bug makes it fly off the rails. But I still wouldn't feel very good about it and would like to do everything I can to avoid it.
Also, when it comes time for you to write life or death code, it would be good if you already knew how to meet the required quality standard.
Actually, there is. Google up "safety-critical software". And maybe "trusted software".
I work in military avionics; every dang line of code in the product, including any libraries we use, is vetted to death. If I were to try to just download a library off the internet and include it in the flight control software, I (A) wouldn't get away with it (B) would probably lose my job and (C) would confuse the hell out of my coworkers who all know I know better than that.
So don't worry someone will include your hastily-developed XML parser in safety-critical software without your knowledge. They won't unless you're willing to prove you've certified it to the level they require. And I promise you, that's not an exercise you'll forget having gone through. ;)
I said there was a process. I didn't say the result was good software. ;)
Keep in mind that things that seem very WTF to you, might seem more plausible when given more details about the subject.
Until we have such a universal standard, it is entirely possible for a bug in your free library to indirectly kill someone, in the course of everyday best practices.
Here is the guidance from the FDA on off-the-shelf software: http://www.fda.gov/MedicalDevices/DeviceRegulationandGuidanc...
> "At the on-board shuttle group, about one-third of the process of writing software happens before anyone writes a line of code. NASA and the Lockheed Martin group agree in the most minute detail about everything the new code is supposed to do -- and they commit that understanding to paper, with the kind of specificity and precision usually found in blueprints. Nothing in the specs is changed without agreement and understanding from both sides. And no coder changes a single line of code without specs carefully outlining the change. Take the upgrade of the software to permit the shuttle to navigate with Global Positioning Satellites, a change that involves just 1.5% of the program, or 6,366 lines of code. The specs for that one change run 2,500 pages, a volume thicker than a phone book. The specs for the current program fill 30 volumes and run 40,000 pages."
From: http://www.fastcompany.com/node/28121/print
If there's no certification for ready-made life critical components, then that means the burden of reviewing, checking and verifying everything in the system is on whoever wants to use it in a life-critical environment.
Intended to be based in Germany in the 60s facing the Russians it wouldn't be powered up for 600 hours, because in 60 hours you would have been overrun.
The Vulcan flight gear famously included hiking boots so the crews could walk to Turkey after completing their mission
Floats are subtle, they can have exactly the kind of nasty little side effects that you overlook during testing and that bite you in a terrible way once you hit production.
Fixed point is the way to go to implement stuff like this, floating point is asking for trouble. Of course floating point is more convenient, but most people that use it don't really understand what's going on under the hood, and most of the time they don't need to.
But bank balances, clock values, GPS coordinates and other values of some importance are best represented in fixed point format. You'll have to do a bit more work when manipulating them but that pays off in reliability.
It didn't, really, and the bug itself is more convoluted than the loss of precision you are describing.
See the references:
http://www.mc.edu/campus/users/travis/syllabi/381/patriot.ht... and http://www.fas.org/spp/starwars/gao/im92026.htm
Inexperienced programmers often regard FP as magical numbers that do everything. It doesn't help that programming languages like JavaScript essentially treat them as such.
Fatal errors are almost never a single thing going wrong unpredictably - they're a chain of events leading to an inexorable conclusion. Root cause - where were these guys?
It is possible to get this right: http://en.wikipedia.org/wiki/Sea_Dart_missile#Gulf_War_.2819...
The reliable (and programmable) digital electronics were originally developed for ICBMs in the 60s, and only after quite a bit of miniaturization were they available for smaller guided missiles, meaning that there wasn't quite as much institutional experience of software engineering among guided missile designers as you'd think.
I think it did amazing job in the 90s against the arguably most technically advanced and innovative military, even more so considering it was pushed past its original specs.
To my eyes, this is a classic training issue. Either the men on the ground were never told to reboot, or they were never told the consequences of failing to do so. Whether that failure was on the manufacturer's part for making crap manuals, or on the Army's part for screwing up the training, I don't know.
http://www.ima.umn.edu/~arnold/disasters/patriot.html
Maybe a better moral is "avoid absolute clocks or counters, if at all possible".
Edited for clarity.
It's very easy to destroy the illusion of a three dimensional graphic on a screen, these are very effective ways of doing so.
A GPS satellite travels much faster than a missile, and is accurate to within meters, so the tolerances are much higher (and the system was designed prior to patriot).
There is a definite 'not invented here' syndrome amongst defense contractors - I doubt they share any information, research or solutions amongst one other, which means the US tax payer foots the bill each time one of these contractors must independently develop and implement a system that has likely already been built in another part of defense.
Yes. Especially because of contractors. A defense contractor is not going to its competitors in order to reveal and exchange know-how. That know-how is a trade secret that gives them competitive advantage. Except that in this case, lack of sharing results in making the same mistake multiple times, which results in loss of human life.
Interestingly lately I have observed that the government is trying to tighten up its defense spending belt and there is a tendency to develop more in-house rather than contract out. In order to gain public support the Congress keeps voting to increase the salary of service personnel. That leaves less money to spend on contracting out. So perhaps we'll see more sharing in the future at least between DoD's own projects.
PS: As you increase accuracy more things become important factors. Large Ship guns actually started to track things like temperature at different altitudes to increase accuracy.
I seem to remember this being the case in my numeric methods course, but maybe that was a different systme.
Bugs are bugs, errors and oversights happen. You have to deal with it, document what happened, and be vigilant to make sure it never happens again.
It's a strong enough expectation that when I had some industrial machines running DOS (albeit on more modern hardware), I added my update, backup & diagnostic scripts to autoexec.bat so that rebooting would fix most of their problems.
It made my life a lot easier, though, because I could update the files and configurations via a master copy on the network, then tell them to reboot everything whenever it was convenient for them (usually between shifts) and the machines would all grab their updates and upload some log files for me to monitor.
I'd like to hear the story of debugging this one. Also how they managed to identify that this incident was caused by that specific bug.
Aside from being a pacifist, this is why a number of engineer friends have stepped out of building defense systems (including missile guidance systems) and into more civilian engineering because the stress and moral burden is just too great.
Also looking at who actually makes the mistake - if someone gives you an update and the system fails, they're at fault. If you give clear instructions for operation and users don't follow it...
Being defenseless for a few minutes every day during a reboot hardly seems like a reasonable fix anyway, especially if it becomes standard procedure that your enemy may learn about.
The fact that the manufacturer did release a patch, rather than a workaround, the day after the accident suggests that this was indeed the safest course of action. It was just taken too late.
Given that it was an issue with a non-terminating binary representation, what would be the way to handle this, without somehow resetting the clock (restart or otherwise)?
Obviously, you could, in a modern system, have more memory and be able to store more bits of the number, but there would still be a limit that you would run into after some amount of time that would cause similar problems.
His objection is that it's too easy for an opponent to defeat the system either by overwhelming it (MIRVs for instance) or by designing the reentry vehicle to make random movements which would make it really difficult for a intercepter to track. Even if the interceptor can hit it, it's more likely to knock the warhead off course instead of destroying it. If a nuclear missile is aimed at NY and an interceptor hits it so that it falls to Philly instead, that's still a net loss.
Saddam inadvertently hit upon the both methods - his engineers tried to improve on the Soviet scud design to give them more range (which they were successful at) but their improvements made the missile more likely to break up on reentry (which presented more targets than the tracking radar anticipated) and the lack of aerodynamics of the resulting pieces (including the warhead) the missiles fall in unpredictable ways which caused tracking problems. An opponent that actually tried to game the system could make his missiles more difficult to hit.
The professor is advocating for boost-phase missile defense since the missile movements are much more predicable.
That airborne laser that's been in development for seemingly forever was pretty much determined to be the best way to intercept. You'll notice that it combines BOTH elements. That 747 is flying within LOS of the missile path while in boost phase, and it's also using the fastest projectile possible.
You'll also note that the Patriot's role is medium tactical air defense as well as theater anti-ballistic missile defense. The second role was basically tacked on, and then later massively expanded (once it became obvious that there aren't many air forces in the world that can actually fight the US).
In the end, I'm sure every general and admiral actually out to improve their warfighting abilities would want both systems. It's all about defense in depth. It's the reason why warships have three different sets of anti-air missiles, while we still have stingers when we have patriot missiles, and why US fighter aircraft still carrying short range missiles and guns.
If you catch something on the launchpad that's a relatively low expense, to catch it i in flight you need to actively defend each and every possible target and a lot of sexy hardware to do so.
I've always thought they should prove that it will work in a realistic physics-based simulation before spending billions on these systems. But they seem to have some value against relatively crude missiles. Plus their value as a bargaining chip, although that depends on how smart the adversary is.
http://portal.acm.org/citation.cfm?id=1251254.1251257 http://www.computer.org/portal/web/csdl/doi/10.1109/HOTOS.20...