An overflow error costing 500M dollars (1996)
around.com
around.com
Part of Feynman's research into the Challenger disaster was finding out how such high numbers at the top level could co-exist with much lower and more realistic reliability estimate by the engineers working on the project.
Anyway, if you like this sort of thing, there's a lifetime supply of reading about horrifyingly expensive software disasters in the famous "RISKS Digest". https://catless.ncl.ac.uk/Risks/
"The Thermocline of Truth." In a big enough company, there is a layer that bad news has trouble getting past.
http://brucefwebster.com/2008/04/15/the-wetware-crisis-the-t...
I don't mean to exaggerate the events. It was a shame nuclear energy suffered such a severe blow. But not as shameful as that parade of engineers who remained stubbornly committed to the consensus risk analysis as events unfolded (and seemingly remain so to this day).
The real question is, why was the diagnostic error message interpreted as flight data?
I recommend reading the original report, which has much more meat: http://sunnyday.mit.edu/nasa-class/Ariane5-report.html
> It was the decision to cease the processor operation which finally proved fatal. Restart is not feasible since attitude is too difficult to re-calculate after a processor shutdown; therefore the Inertial Reference System becomes useless. The reason behind this drastic action lies in the culture within the Ariane programme of only addressing random hardware failures. From this point of view exception - or error - handling mechanisms are designed for a random hardware failure which can quite rationally be handled by a backup system.
> Although the failure was due to a systematic software design error, mechanisms can be introduced to mitigate this type of problem. For example the computers within the SRIs could have continued to provide their best estimates of the required attitude information. There is reason for concern that a software exception should be allowed, or even required, to cause a processor to halt while handling mission-critical equipment. Indeed, the loss of a proper software function is hazardous because the same software runs in both SRI units. In the case of Ariane 501, this resulted in the switch-off of two still healthy critical units of equipment.
> The original requirement acccounting for the continued operation of the alignment software after lift-off was brought forward more than 10 years ago for the earlier models of Ariane, in order to cope with the rather unlikely event of a hold in the count-down e.g. between - 9 seconds, when flight mode starts in the SRI of Ariane 4, and - 5 seconds when certain events are initiated in the launcher which take several hours to reset. The period selected for this continued alignment operation, 50 seconds after the start of flight mode, was based on the time needed for the ground equipment to resume full control of the launcher in the event of a hold.
> This special feature made it possible with the earlier versions of Ariane, to restart the count- down without waiting for normal alignment, which takes 45 minutes or more, so that a short launch window could still be used. In fact, this feature was used once, in 1989 on Flight 33.
> The same requirement does not apply to Ariane 5, which has a different preparation sequence and it was maintained for commonality reasons, presumably based on the view that, unless proven necessary, it was not wise to make changes in software which worked well on Ariane 4.
> Even in those cases where the requirement is found to be still valid, it is questionable for the alignment function to be operating after the launcher has lifted off. Alignment of mechanical and laser strap-down platforms involves complex mathematical filter functions to properly align the x-axis to the gravity axis and to find north direction from Earth rotation sensing. The assumption of preflight alignment is that the launcher is positioned at a known and fixed position. Therefore, the alignment function is totally disrupted when performed during flight, because the measured movements of the launcher are interpreted as sensor offsets and other coefficients characterising sensor behaviour.
It's been a while but I wrote about this here: https://blog.bugsnag.com/bug-day-ariane-5-disaster/
That does not actually follow. It would not have crashed due to this specific issue, but it might have crashed from some other issue.
That got me to thinking, could 3 different teams have separately built and programmed 3 different units (separate software and hardware), and then had a voting type system so that, if say 2 units are saying "keep going straight" and the third is saying "turn left", then it knows to keep going straight? Is something like this ever done in systems that require extremely high reliability?
> why the hell is the video called "2002" is this is clearly the footage from 1996 ? you even wrote it in the description ! moron
You may also be interested in section 2.3 which describes the factors that led to simulations failing to discover this failure mode
[1] https://web.archive.org/web/19970125194224/http://www.esrin....