Infamous Software Bugs
bbvaopenmind.com
bbvaopenmind.com
The mailing list and digest just celebrated it's 30th birthday in August. There's an incredible wealth of history here regarding bugs, related issues, and the overall risks of our dependency on computers in society.
RISKS broke most of the major stories of the day including the Morris Worm, the AT&T Long Distance collapse, the Mars Climate Orbiter unit error, the Mars Pathfinder priority lockup, the Therac-25, the first email spam, the Pentium FDIV bug, and thousands of other interesting, amusing, and/or scary bugs.
http://catless.ncl.ac.uk/Risks/7.69.html#subj1
Warning - this one email will lead you to multiple rabbit holes (Stoll, The Cuckoo's Egg, Robert Morris, etc).
As others have noted, the big miss is the Therac-25 bug, which is pretty commonly taught in Software Engineering/Computer Science classes the example of how intangible software can kill people.
This is one of those ideas that sounds a lot easier than it likely is.
You can do real time dose monitoring using diodes (see [1]). The diodes are quite small and don't generate huge amounts of scatter.
It is quite common to use TLD's[2] to monitor patient dose which is then analyzed offline to give dose estimates after the fact.
edit: that said, I don't think anybody was doing real-time in-vivo measurements during the era of the Therac incident.
[1] https://www.aapm.org/pubs/reports/RPT_87.pdf
[2] TLD=thermoluminescent dosimeter, little hunks of metal whose electrons get stuck in an excited state when exposed to radiation, you then heat them up and measure the light they give off as they relax back to ground state. The amount of light they give off can then be correlated with the dose the patient received.
Obviously, it's too late at that point to prevent the beam from hitting the patient, but you'll know that something went wrong and can lock the system until the problem is found.
[1] http://tommytoy.typepad.com/.a/6a0133f3a4072c970b0147e2ed7f8...
As an aside, TLDs and IVD systems are falling out of favor for patient dose monitoring because AAPM TG 62 (referenced in the link in your other post) is not directly applicable to IMRT and VMAT modalities, which are pretty common these days.
[1]http://www.sunnuclear.com/medphys/patientqa/epidose/epidose....
[2]http://www.sunnuclear.com/snc_site/solutions/patientqa/perfr...
p.s. If you are employed in Med Phys allow me to make a quick plug for my open source Med Phys QA database project: http://qatrackplus.com/
That's how I'd do it from my understanding of the usage practices today.
http://dealbook.nytimes.com/2012/08/02/knight-capital-says-t...
almost ;-)
[1]: http://www.ccnr.org/fatal_dose.html [2]: http://www.ibiblio.org/harris/500milemail.html
http://wiretap.area.com/Gopher/Library/Techdoc/Lore/rumor.ne...
Software Horror Stories : http://www.cs.tau.ac.il/~nachumd/verify/horror.html
It's bit hard to digest. ( Although just checked wikipedia, it also says so ) How can a high performance organisation like NASA could make such a simple yet fatal mistake ?
Wikipedia page of Mars Climate Orbiter says that NASA was informed about this discrepancy by two people, but the "concerns" were dismissed.
What am I getting wrong here ? These are not the "concerns" you simply dismiss in a space mission. Could there be another story to this ?
…At least, that's the official story. I recall there being a lot of conspiracy theory-like buzz at the time from people who also couldn't believe NASA could make such a stupid error like that. It does make you wonder.
The report does go ahead and state all sorts of organizational (and otherwise 'soft' issues) that contributed to the end failure.
The report notes that earlier deviations between measured and modeled results were noted, however, they were hampered by limited data (in the sense that they couldn't measure what they wanted). It is implied (though not stated) in the report that in the absence of appropriate data, the operations navigation team attempted to contain/mitigate the deviations instead of 'solving' it.
The report also notes substantial organizational issues. Different navigation teams were used in development and operations, and there were insufficient knowledge transfer during hand-off that hampered the operations navigation team ability to notice these issues. Communications between the main operations team and the ops nav team were not effective. They were apparently spatially separate teams. In addition, model-measurement conflicts which were brought up were solved via e-mail instead of over formal processes. The report suggests that systemic use of formal processes may have allowed teams to uncover the problem earlier in time.
And of course, the report also states that insufficient verification/validation of the supplied software was not completed. The entire section on verification/validation (MCO Contributing Cause No. 8) is just a giant cringe fest.
The implication is that the MCO project was just... not run well.
So had a look at the report.
There was one more problem actually. This machine, the MCO, had asymmetrical solar panels which would cause solar pressure ( force by sunlight ) to create a very mild spin ( angular momentum ). Now this angular momentum had to be desaturated time to time in order to keep this machine stable. Now, one module called SM_FORCES calculate this adjustment and feeds to AMD ( Angular Momentum Desaturation ). Now, this SM_FORCES & AMD uses different unit system, which was ignored by whoever wrote this connecting piece of program. Due to this error desaturation was not enough ( or more ) and it kept building over the period of 9 months.
Now, I notice that NASA has a separate team to navigate this machine to mars. There data showed this angular momentum adjustment event occurred 10-15 times more than expected. It was like a man walking with one leg shorter than another. It's a 9 months journey from mars to earth. They must have seen the first sign to inconsistency with in first few weeks only, just guessing though.
In this report, out of 8 possible contributing causes, at-least 3 are attributed to navigation team. I think success of such mission depends not only on meticulous planning but also on thinking on the feet ability of the team. ( Any Apollo 13 fans? :) )
"e. Two weeks before the incident, Army officials received Israeli data indicating some loss in accuracy after the system had been running for 8 consecutive hours. Consequently, Army officials modified the software to improve the system's accuracy. However, the modified software did not reach Dhahran until February 26, 1991--the day after the Scud incident."
There are two ways to run a PHP app on nanobox. You can eitherconfigure the generic ruby engine, or use a framework specific engine.
http://ntrs.nasa.gov/archive/nasa/casi.ntrs.nasa.gov/1991000...
The story is normie clickbait anyway, and most of the bugs aren't mismatches between the source code and the (possibly non-existent) unit testing infrastructure, they're just cultural examples of blaming the lowest social status individual involved, that usually being a programmer. There was a programmer involved, someone in management screwed up and doesn't feel like taking the blame, therefore its the programmer's fault. In the olden days they'd just have blamed the closest (insert ethnic group here) or (insert religious group here), nothing really to be proud of.
this page is bullshit.