The Therac-25 Incident
thedailywtf.com
thedailywtf.com
I am a physician who did a computer science degree before medical school. I frequently use the Therac-25 incident as an example of why we need dual experts who are trained in both fields. I must add two small points to this fantastic summary.
1. The shadow of the Therac-25 is much longer than those who remember it. In my opinion, this incident set medical informatics back 20 years. Throughout the 80s and 90s there was just a feeling in medicine that computers were dangerous, even if the individual physicians didn't know why. This is why, when I was a resident in 2002-2006 we still were writing all of our orders and notes on paper. It wasn't until the US federal government slammed down the hammer in the mid 2000's and said no payment unless you adopt electronic health records, that computers made real inroads into clinical medicine.
2. The medical profession, and the government agencies that regulate it, are accustomed to risk and have systems to manage it. The problem is that classical medicine is tuned to "continuous risks." If the Risk of 100 mg of aspirin is "1 risk unit" and the risk of 200 mg of aspirin is "2 risk units" then the risk of 150 mg of aspirin is strongly likely to be between 1 and 2, and it definitely won't be 1,000,000. The mechanisms we use to regulate medicine, with dosing trials, and pharmacokinetic studies, and so forth are based on this assumption that both benefit and harm are continuous functions of prescribed dose, and the physician's job is to find the sweet spot between them.
When you let a computer handle a treatment you are exposed to a completely different kind of risk. Computers are inherently binary machines that we sometimes make simulate continuous functions. Because computers are binary, there is a potential for corner cases that expose erratic, and as this case shows, potentially fatal behavior. This is not new to computer science, but it is very foreign to medicine. Because of this, medicine has a built in blind spot in evaluating computer technology.
https://www.theguardian.com/world/2020/oct/26/tens-of-thousa...
https://globalnews.ca/news/6311853/lifelabs-data-hack-what-t...
1) The software influences the medical workflow and becomes a major distraction to visits. What was completely analog and free-form is now binned, discrete, and made more complex.
2) Desired outcome affects what and how things are recorded, instead of vice-versa. Staff learn what they have to do to get the orders or prescription or billing output that they need, which often doesn't line up with the actual diagnosis.
3) It reduces productivity, causes physician burnout, and put most small practices out of business. (Any practice not large enough to self-host records hire and have full time IT/software staff were basically strong armed into selling out to hospitals or other large org.)
4) It tracks both efficiency and patient satisfaction in the same system as the records and billing, which leads to some pretty perverse incentives on multiple levels. (As much heat as the drug companies over the opiod crisis, I'd argue doctors worried about dinging their patient satisfaction scores were just as responsible, by being afraid to tell too many patients "no".)
It goes on... Just a huge cluster of poorly thought out unintended consequences.
And that's just the medical side of things. The technical, financial, and legal aspects all have similar issues.
This sentence makes zero sense. Limiting choices inherently makes things LESS complex. That's why we use frameworks for decision making and risk assessment, rather than just doing everything "analog and free form"; it's the same reason we do "structured training" for complex and difficult tasks, rather than just let people try to figure them out. It's just a completely wrong and backwards statement.
"Do you want me to kick you in the face, or the groin?" is a substantially more complex choice than "Do you want me to kick you in the face, or the groin, or not at all?"
Sometimes, limited choices require you to shoehorn something into one of those available choices when it's not actually appropriate. Such is the case with medical record systems at times.
To address your points
Half of this is a good thing. I want my doctor to follow the proper process and checklist everytime, no matter what. There are many one in a million cases that have the same symptoms as the common thing they see daily. The process is how you catch them and treatment. The doctors office is no place for creativity until everything else has been ruled out.
That doesn't mean there isn't room to improve the user interface. However doctors are in the wrong to be so technology backwards
And frankly, this type of response (the software knows better than the doctor/human) is (again) a root cause of problems like the Therac-25 incedent.
The system needs to conform to the doctors, not vice-versa.
User interaction studies have made great progress. A lot of software ignores that. Likewise we know a lot of ways to write high quality software that are ignored in your typical web app.
However we also know that doctors are human and they fail often. while computer systems do fail, those failures are much easier to fix once and for all.
Continuity of risk, change, incentives, etc. lend themselves to far easier analysis and confidence in outcomes. And higher degrees of continuity as well as lower values of change only make that analysis easier. Of course it's a trade-off: a flat line is the easiest thing to analyze, but also the least useful thing.
In many ways I view the core enterprise of planning as an exercise in trying to smooth out discontinuous jumps (and their analogues in higher degree derivatives) to the best of one's ability, especially if they exist naturally (e.g. your system's objective response may be continuous, but its interpretation by humans is discontinuous, how are you going to compensate to try to regain as much continuity as possible?).
If this was completely analog, mechanical tool, the problem could still have existed. For example, you could have made a mistake using various knobs and switches.
And in this case the solution was to physically design the device so that it is not possible to put it in an incorrect state. Thus it did not matter much if there was a software error as whatever software did would not make it possible to put the beam and the metal shield in an incorrect state.
Using physical safeties is a common strategy. For example, my Instant Pot has multiple physical safeties built in.
For example:
* even if the controlling software or sensor fails, there is overpressure valve that will not let the pressure rise over certain value.
* there is a single piece of metal that both blocks gasses escaping from the pot AND blocks ability to open the pot. If the piece of metal is not in place, there is a hole in the pot and pressure cannot rise. If the piece of metal is in place so the hole is closed, it blocks possibility of opening the pot and creating an explosion.
* there is a specially designed guide that makes it impossible to close the device only partly
* the device is designed so that the weakest element keeping pressure is the seal. If the pressure rises, instead of the pot blowing up the seal will deform, be blown off and let the steam escape in more or less controlled manner.
* there is a bimetal safety device that will turn off the heater if temperature rises too much,
and so on.
You see, none of these features relies on software. There are software safeties but they are redundant in that device does not rely solely on them.
If the company producing Instant Pot can show this kind of safety-conscious design I am sure it can also be applied to dangerous medical devices.
2019 https://news.ycombinator.com/item?id=21679287
2018 https://news.ycombinator.com/item?id=17740292
2016 https://news.ycombinator.com/item?id=12201147
2015 https://news.ycombinator.com/item?id=9643054
2014 https://news.ycombinator.com/item?id=7257005
2010 https://news.ycombinator.com/item?id=1143776
Others?
For example, the issue which caused this in 2007:
https://www.heraldtribune.com/article/LK/20100124/News/60520...
Or the process issues which caused this in 2001:
https://www.fda.gov/radiation-emitting-products/alerts-and-n...
> As Scott Jerome-Parks lay dying, he clung to this wish: that his fatal radiation overdose -- which left him deaf, struggling to see, unable to swallow, burned, with his teeth falling out, with ulcers in his mouth and throat, nauseated, in severe pain and finally unable to breathe -- be studied and talked about publicly so that others might not have to live his nightmare.
which is, once you start blowing off testing, having exceptions to avoid every fix, you rarely if ever leave that mode and end up with a stressed support staff and users in the field who think your coders are incompetent.
Are we actually not as bad at mitigating risks when needed as the rest of the industry would led us to believe, or are those incidents just not covered?
I do not mean this as a judgement on those who do work on systems that can physically harm people. Obviously, we need good engineers to design potentially dangerous systems. It is just how I realized I really don't have the character to do it.
There are plenty of other high stakes software that involve human lives (Uber self driving cars, SpaceX, Patriot missiles) and many of them completely scare me and morally frustrate me as well to the point where I would not like to work on one, but I totally understand if you have a personal profile that is different than mine.
Automotive software engineer here: I've asked the same in interviews.
"We work on multi-ton machines that can kill people" is a frequently uttered statement at work.
On the other hand, I'm not entirely sure it's appropriate to be comfortable with that kind of position in any case.
In industrial non software settings this is not as rare as you'd think, if you research the tower crane, truck crane and rigging industry. Lots of things can kill people. The important part is that we need the appropriate 'belt and suspenders' safety checks and engineering practices to prevent them from doing so.
randomly chosen industrial example:
Not sure if this is still covered today.
This makes me think - there was only one developer there, I guess, who was doing everything in assembly. This software, and the process to produce it, must have been designed in the early days of their devices, when there would be expected to be hardware interlocks to prevent any of the really bad failure modes. I bet they never did change much of the software, or their procedures for developing, testing, qualifying, and releasing it in light of the change from relying on hardware interlocks to the quality of the software being the only thing preventing something terrible from happening.
> Related problems were found in the Therac-20 software. These were not recognized until after the Therac-25 accidents because the Therac-20 included hardware safety interlocks and thus no injuries resulted.
The safety fuses were occasionally blowing during the operation of Therac-20, but nobody asked why.
Have you tried turning it off and on again?
I always liked the testing philosophy institutes like for instance Underwriter Laboratories had: Your product will fail. This is stated as fact and is not debatable. What kind of fail safes and protection have you made so that when it fails (and it will) it cannot not harm the patient?
Single-point failure analysis techniques like fault trees, event trees, and especially the tabular "failure modes and effects analysis" (FMEA) are so powerful, especially for safety-critical hardware, that when people learn them they want to apply them to everything, including software.
However, FMEA techniques actually have been found to not apply well to software below about the block diagram level. They don't find bugs that would not be found by other methods (static analysis, code review, requirements analysis etc) and they're extremely time and labor intensive. Here's an NRC report that goes into some detail: https://www.nrc.gov/reading-rm/doc-collections/nuregs/agreem...
Even more interesting! Thank you for this link. I appreciate it. Never too old to learn. :)
I'm not sure if the settings were locked as a default by the manufacturer, or a protected setting that no one knew how to change, or simply locked by management.
Anyhow, from mild atrial fibrillation to irregular rhythm, the alarm was constantly beeping, which was to be expected, considering the acuity of our patients. Nurses and MD became really annoyed by this.
Eventually, one of the staff covered the speaker from inside, to muffle the almost constant warnings. The staff repeatedly asked management to find a solution to this, and the associated risk.
One day, a patient had a life threatening cardiac incident. The alarm was muffled enough that no one nearby could hear a thing.
I agree with the sentiment that they took the software for granted. I get the feeling that happens in a lot of settings, most of them less life-threatening than this one. I've come across it myself too, in finance. Somehow someone decides they have invented a brilliant money-making strategy, if they could only get the coders to implement it properly. Of course the coders come back to ask questions, and then depending on the environment it plays out to a resolution. I get the feeling the same thing happened here. Some scientist said "hey all it needs is to send this beam into the patient" and assumed their description was the only level of abstraction that needed to be understood.
> I am a physician who did a computer science degree before medical school. I frequently use the Therac-25 incident as an example of why we need dual experts who are trained in both fields. I must add two small points to this fantastic summary.
> 1. The shadow of the Therac-25 is much longer than those who remember it. In my opinion, this incident set medical informatics back 20 years. Throughout the 80s and 90s there was just a feeling in medicine that computers were dangerous, even if the individual physicians didn't know why. This is why, when I was a resident in 2002-2006 we still were writing all of our orders and notes on paper. It wasn't until the US federal government slammed down the hammer in the mid 2000's and said no payment unless you adopt electronic health records, that computers made real inroads into clinical medicine.
> 2. The medical profession, and the government agencies that regulate it, are accustomed to risk and have systems to manage it. The problem is that classical medicine is tuned to "continuous risks." If the Risk of 100 mg of aspirin is "1 risk unit" and the risk of 200 mg of aspirin is "2 risk units" then the risk of 150 mg of aspirin is strongly likely to be between 1 and 2, and it definitely won't be 1,000,000. The mechanisms we use to regulate medicine, with dosing trials, and pharmacokinetic studies, and so forth are based on this assumption that both benefit and harm are continuous functions of prescribed dose, and the physician's job is to find the sweet spot between them.
> When you let a computer handle a treatment you are exposed to a completely different kind of risk. Computers are inherently binary machines that we sometimes make simulate continuous functions. Because computers are binary, there is a potential for corner cases that expose erratic, and as this case shows, potentially fatal behavior. This is not new to computer science, but it is very foreign to medicine. Because of this, medicine has a built in blind spot in evaluating computer technology.
Consider something like a surgeon nicking and artery while performing some routine surgery, the patient not responding normally to anesthesia or the anesthetist not getting the mixture right and the patient not coming back the way they went in. Or that subset of patients that have poor responses to a vaccine.
Everybody likes to think of the world as a linear system, but it's not.
In other words: tangible objects usually correspond to what we see; in software, you have no way necessarily of knowing if the UI/interface is outright lying to you. It could be doing anything internally, and a single flipped bit deep in some subroutine could cause death.
Highly recommend all of his videos, actually.
https://direct.mit.edu/books/book/2908/Engineering-a-Safer-W...
Some questions in my mind while reading this article (but I couldn't find them quickly in a search): Who were the executives that were running the company? Sounds like something that should be taught to MBA as well as CS students. Further, the AECL was a crown corporation of the Canadian government. Who was the minister and bureaucrats in charge of the department? What role did they have in solving or covering up the issue?
I know it well from it being the first and main case study in my software testing class as an undergraduate CS major in Washington DC in 1999.
It will never not be interesting.
It’s because it can be the difference between “Shoot” and “Don’t shoot.”
I feel that lot of manual errors and UI errors can be mitigated by careful design. Punishing natural careless errors is not very effective, as recent research shows [1]:
[1] https://www.jstage.jst.go.jp/article/jjsca/32/7/32_954/_arti...
One another example that comes to my mind is the Toyota "unintended acceleration" incident. Or the "Mars Climate Orbiter" incident.
We both posted the EXACT same video at the exact same time!
Plainly Difficult is an awesome channel.
I am not working on safety-critical systems but still, the failure of the type of backend systems I work on can bring huge damage to the company.
You can come out with a set of rules that can really improve the reliability of your application.
* Ensure the system always moves between known, valid states. Errors can still be valid states if you plan for them correctly.
I like to compare this to logical induction: if you start with correct state and on every operation ensure that if it executes on correct state it must result in correct state, then you guarantee that the system is always in correct state.
* Ensure incorrect operations are impossible to execute.
As an example, if you want to guard against large data loss, you can choose to not expose ability to remove data. In my current project we took away everybody's write access to the database, we removed all functions that can remove the data from the database and exchanged them with functions that either mask data (setting flags so that application ignores it) or move to archive. If you need to run an operation on database, you need to submit a pull request and go through Code Review. There is additional review before merge where the team goes through checklist to verify the change meets certain standards.
* Ensure anything the client can throw at the application does not have a chance to escalate to other transactions.
All inputs are limited and validated mercilessly, all queries are constructed so that it is possible to predict resources necessary to execute the query. Every type of query has a limit on how many queries of that type can be executed in parallel or in a unit of time and how much results it can return.
All functionality is constructed in such a way that the performance per transaction does not degrade when load increases (for example by amortizing costs and batching work). Then every functionality is load tested and based on testing results it is certified with maximum load that it can process. Everything over that limit is preemptively declined.
* If an algorithm or a solution cannot be guaranteed to work correctly it cannot be used.
For example, if library cannot be guaranteed to use at most X amount of memory for the operation, then it is not suitable.
* Every single failure of the system is investigated.
No system is without faults. The only way to build reliable system is to not ignore errors and always try to figure out underlying failure in the process that let the fault happen.
Most projects deal with errors only after the error is urgent enough or only to fix the immediate cause of the error, rather than the failure in the process.
If you are interested in very reliable software, when you have it fail for whatever reason you need to look at all stages of your process and identify all places that were supposed to prevent the problem but did not.