Killed by a Machine: The Therac-25
hackaday.com
hackaday.com
I asked him how they could possibly use such machines on humans, given what I just watched it do to a 1kg plate of acrylic. He told me that they hit the plate with way more energy than they ever would a human. That prompted my followup question: "uncle, what happens if you accidentally hit the wrong button?" He told me that accidents like that used to happen, but that the machines they used had special computers to keep the patient safe even if he made a mistake. "But uncle, what happens if the computer makes a mistake!?" I had no idea what a bug, or for that matter code was. He didn't have a good answer beyond "the computer can't make mistakes like we do." Having played with computers enough by then to know that his statement wasn't entirely true, I ended that outing wanting to know more. I started obsessing over what would happen to people if computers controlling things like that linear accelerator, or even the elevator in my dad's office building, made a mistake.
Incidentally, my uncle was the one who got me interested in science, and that trip to the hospital got me using computers for something other than games. Fast forward 24 years and...well...part of what I do is work on provably correct systems.
I work in software engineering, so I'm exposed daily to broken software controls, and I'm gradually becoming more of a "grumpy old man" longing for the good old days when machines (be it cameras, cars, watches, medical equipment or anything else, really) could be "debugged" by following levers, wires, physical stops and hoses. I feel much more comfortable with that than with complex computer systems.
I would like to know if my fears can be founded in actual truths, or if I'm just riling myself up over nothing.
I spent over 10 years writing control software for medical devices. In classifying the criticality of a device, the FDA from the beginning put any device with software into a higher category. They have since refined their classifications, so you probably won't find that guidance on their site anymore, but 15 years ago it was pretty eye opening to me that two devices of the same type, one with and one without software, could be automatically in different safety categories.
So, no, it's not just you :-)
I think it's in regard to risk controls, but I'm not certain of that.
Edit: Just looked it up: "If the HAZARD could arise from the failure of the SOFTWARE SYSTEM to behave as specified, the probability of such failure shall be assumed to be 100%"
Thank you for a helpful observation. It had not occurred to me.
The "trick", or rather, the true test of a design when it comes to this is judiciously balancing the advantages that each of these categories bring you. You generally want to use "hardware" means in order to nullify the important risks of programmable logic (i.e. the major consequences of a blunder in the logic that was programmed on the device), and use programmable logic for those sections that require flexibility, ease of programming, potential fixes or enhancements and so on.
Real-life example from a device I worked on: the linear actuator (that was essentially sticking needles into people's brains) had a physical barrier past the safe distance limit. Literally a big hunk of metal that the actuator could not be moved past.
Of course, there was a limit in software as well (the software would refuse to move the motor past the safe limit). However, with the physical barrier in place, it protected the motor more than anything else.
It's worth noting that design decisions in these fields are done not so much based on the amount of brokenness in existing implementations (to put it bluntly, there's about as much broken hardware as there is broken software), but based on the amount of risk that a solution introduces. Many regulating bodies (e.g. FDA) will classify your device in a higher risk class if there's programmable logic in it, simply because software tends to be harder to write and test reliably (not only because of bad programmers, but also because of a lack of standards, or at least consensus and metrics on testing).
Edit: I can't quote numbers right now; there is data that shows the risks involved in purely software-based approaches, and it's easy to see why even in the example above. A software-enforced limit on the distance of movement can fail due to a bunch of reasons, not all of them bugs. Bugs are one thing, but the system could also fail due to a broken connection between the motor driver and the CPU, due to a bug in the hardware implementation of the motor driver itself, due to glitches on the bus and so on.
There are software mitigations for all of these cases, too (e.g. you continuously monitor the position of the motor; if the motor still moves after you tried to stop it, you reset the system, and the hardware is wired so that all power to the motor is cut when you come out of reset), but nothing is as efficient as making sure that the motor just won't be able to move the load past a certain distance by placing a physical barrier in its way. By an uncanny chain of events, maybe all the software mitigations can fail; nothing, however, can make a big hunk of metal disappear into thin air.
Obviously, you can't go down the rabbithole forever, and this is a bit tongue-in-cheek, but IIRC similar circumstances have occurred.
The San Salvador medical irradiation facility accident[1] is largely a tale of gradual failure of safety features and interlocks, combined with the misunderstanding that they were still adequate or safe, until they weren't.
[1] http://www-pub.iaea.org/MTCD/publications/PDF/Pub847_web.pdf
That wouldn't be seen as the fault of the device. Software bugs will still show up when the product is used as directed. Ideally, though, the device wouldn't be able to be used without the physical stop.
Furthermore, the design process is shaped so that it leads to such designs. If the first iteration of the design had included a removable physical stop, it would have been caught in the risk analysis. Most mission-critical fields have standards that specify this sort of stuff (IEC 60601 for medical devices, DO-254 and DO-178C for avionics etc.)
When designing this kind of devices, whether it's the fault of the device or not is not too relevant in general. You have to mitigate any kind of unacceptable risk (i.e. things that lead to injury or death).
There are certain common-sense exceptions here. For instance, the device isn't expected to operate properly outside its specified operation conditions, but you have to clearly state what those conditions are (altitude, temperature, humidity, presence or absence of liquids etc.) and put warning labels regarding their breach in the manual). Similarly, the level of mitigation is often specified by standards. E.g. for medical devices, IEC 60601 specifies insulation requirements that will protect against the kind of shocks you could get from a faulty network, but not if the device is struck by lightning or strapped to an electric chair. IEC 61010 (Safety requirements for electrical equipment for measurement, control and laboratory use) similarly includes provisions for the kind of protection you would need in equipment that falls under a specific type of use (e.g. here on Earth, not up there in space).
Chainsaws cut wood, I think you meant an angle grinder or a blow-torch.
Apparently the organization continues to exist under the name "Nordion".
Ah, so if I click these angular nav-pills very quickly in series without waiting for the page to load, something unexpected might happen?
What is annoying in one context may be deadly in another.
----
If the "Sentinel Event" policy had been in place at the time, perhaps these deaths would have been prevented.
A sentinel event is any event that either leads to death or serious permanent injury -OR- could have lead to death or serious permanent injury.
1. https://www.jointcommission.org/sentinel_event_policy_and_pr...
Yet, when a poorly designed self-driving car kills people, many rush to point out (correctly or not) all the lives that were saved by the technology.
The valley seems to have an entirely different level of regard for the consequences of their work then engineers and doctors do.
Also, it's easy to imagine self driving cars that are pretty good, but still occasionally kill people, yet at a rate much lower than human drivers.
For the Therac-25 on the other hand, it's not so complicated. Why it failed and how to fix it is already well known.
It's easy to imagine a medical device that's pretty good, but still occasionally kills people because of flaws in its design. The image is fairly terrifying, actually.
We could also say the same about the fatal Tesla accident - we know why it failed, and how to fix it... And observe how quickly large parts of the tech community (I would expect no less of the vendor) was to blame the human.
As an engineering student it was meant to imprint one thing upon us, but talk about a dark introduction to the word 'responsibility'...
In the case of the fountain, several couples had had a night out and decided to splash in the fountain. The first two knocked a power conduit for one of the pumps loose, electrocuting themselves. The others died as they each entered the fountain to rescue the others. The system lacked a GFCI protector.
While code we write doesn't always have such dire consequences, it was an eye opener to me as a freshman in college. It definitely made designing software/hardware to never fail have a much higher priority on my list of priorities.
The fact that a simple error could cause someone to lose a limb or die does wonders for your focus.
This is actually quite a tricky thing to do right because to be able to jog the machine out of a shut-down like that you have to re-enable it in a potentially un-safe situation. For each and every little challenge like that we found a good solution but some of those were real head-scratchers.
A really nasty one that I recall was that when you power up a bunch of latches they can come up in an undefined state, so the decision was made to include a detector for that undefined state which first would have to be cleared before the output of the latches was allowed to influence the motors.
This worked well in practice but given the restrictions of the machine this was all done on it took a bit of thinking, the solution we settled on was a magic sequence output on the parallel port indicating the system had successfully reset after which the relay powering the motor drivers would engage. So until that relay triggered everything else was ignored.
This was already important enough with stepper drivers, but once we switched to servos for some of the more demanding applications it became crucial to safe operation that the drivers would never ever be energized with faulty inputs. A servo driving a ball-screw will happily wreck itself, the machine it is bolted on to and anything standing in-between (including the operator) if it suddenly gets driven to -10 or +10 V and naturally, you'd always get one of those two, never the safe '0'.
Separation of duties is a good principle, and wherever possible you should use it.
https://en.wikipedia.org/wiki/Separation_of_duties
But in the case of a single tech guy in a company there isn't a whole lot you can do in that direction, so it is best to clearly assign ownership and make sure that the people involved realize full well the consequences of a fuck-up.
They wrote code that depended on hardware controls, didn't document their reliance on the hardware controls and killed a bunch of people. DOCUMENT ALL YOUR DEPENDANCES!!!
...similar to Toyota's electronic throttle control system design.
Also, if you rely on hardware features for safety, it's still ok/good to design the software as if it didn't depend on those features whereever possible.
At some level you can no longer abstract away the situation and the system has to perform as a system. You can write soft controls all you want but if hard controls are present in the spec and something that's not likely to change there becomes a point where chasing down every little problem and checking for every possible error is no longer cost effective. Nobody cares if you write Mars rover tier code for a bulldozer and spending the resources doing is wasteful if your competitors aren't also spending the resources. Obviously you can write total crap that's outside the acceptable range toward the other end.
If you're designing software to control a widget that moves and does stuff has a hard switch to prevent your widget's equivalent of an out of battery detonation you have to strike a balance between relying on it and introducing complexity from soft controls. Every aspect of the system is involved in making that determination.
If you've got a hard switch in most cases you may as well write a five line timeout controlled while loop that tries to perform the action and waits and tries again if there's no feedback indicating the action was performed. As long as nobody removes that switch and the code's dependency on it is very obviously documented then that simple unsafe code is probably better than more complex code that performs a redundant (redundant because you have a hard switch) check because the more complex code has more going on.
Still doesn't make a difference if there is no place for the documents to live.