I think the finest example I have seen was an asset tracking system in the defence sector. Independent auditors came in to make sure that secure assets were shredded. They found stuff that apparently wasn’t. This spawned a large panic. I sat through an hour long finger pointing session with the auditors, disposal staff, programme manager and engineering manager all who had different stories and none of them thought it was a software issue. They had got to the point of calling each other liars and swearing.
The problem which I found in two minutes flat? Well the thing that returned the state of the asset did a SQL join on the history and returned the top row’s state without sorting by date. By coincidence this had probably looked like it worked in dev. If you closed and opened the screen a few times it would return a different state because of non determinism in the query. Adding one order by and it was resolved and the software was the problem and all assets were disposed of. Then I was tasked to write test cases for it and found a hundred other nasty things.
fortunately it is not as often as some think it is, but delegation of tasks still does require those delegating the work to verify it is done correctly, whether by machine or man. in all cases any project should have documented means of verification anyone can act upon.
what tends to break this is having too many responsible parties to where no one is responsible for any single point of failure and this insulates the the whole. (the old too many cooks)
The thing is, a computer doesn't generally make mistakes. Not unless some hardware failure happens, or a photon flips some bit in an unfortunate moment. It's always humans who make mistakes - computers are just executing them blindly. They're very good and fast at it.
Now I know it may sound trite, but I don't think it is. The problem isn't people thinking computers can't make mistakes. The real problem is that people are thinking that the output from the computer is something else than it is. Taking m0xte's example[0]:
"Well the thing that returned the state of the asset did a SQL join on the history and returned the top row’s state without sorting by date."
The mistake number one is, the function that did this was probably called "getCurrentAssetState", or something else which implied it's getting the current state of the asset. Some programmer made a mistake here. Then that description was taken at face value and the problem traveled all the way to end-user level, probably to some label with "current status" on it.
But that's not the only type of mistakes that can happen. Another possible mistake: the data in the database was wrong, or stale. Another: eventual consistency. Another: configuration error. Etc. And regardless of that, the computer was always doing the thing it was told to, without any error: SELECT state FROM asset JOIN asset_history ON asset.id = asset_history.asset_id LIMIT 1;.
My point is: once the discussion stops being about whether or not a computer can make an error (generally, it can't), people can start to appreciate that they're dealing with large systems designed by humans - and both in the large and in the small, these systems do something resembling what they were designed for, but never exactly that.
--
HAL. Hello, HAL, do you read me? Hello, HAL, do you read me? Do you read me, HAL? Do you read me, HAL? Hello, HAL, do you read me? Hello, HAL, do you read me? Do you read me, HAL?
HAL: Affirmative, Dave. I read you.
Dave: Open the pod bay doors, HAL.
HAL: I'm sorry, Dave. I'm afraid I can't do that.
Dave: What's the problem?
HAL: I think you know what the problem is just as well as I do.
Dave: What are you talking about, HAL?
HAL: This mission is too important for me to allow you to jeopardize it.
In the end no matter what the computer does, it was programmed by humans.
The computer is always right, but the intention might be wrong. This gets scarier when it is a neural network or AI decision that is not known even to the programmer, only comes up in data the algorithm understands. Lots of edge cases.
It would suck to be stuck in a bad decision/error/bad interpretation of data and nothing you can do about it. No customer support to help you, just the lifeless Borg deciding your fate.
HAL: Well, I don't think there is any question about it. It can only be attributable to human error. This sort of thing has cropped up before, and it has always been due to human error.
Frank: Listen HAL. There has never been any instance at all of a computer error occurring in the 9000 series, has there?
HAL: None whatsoever, Frank. The 9000 series has a perfect operational record.
Frank: Well of course I know all the wonderful achievements of the 9000 series, but, uh, are you certain there has never been any case of even the most insignificant computer error?
HAL: None whatsoever, Frank. Quite honestly, I wouldn't worry myself about that.
Dave: Well, I'm sure you're right, HAL. Uhm, fine, thanks very much.
We should take decisions by systems as we do foreign policy: "Trust, but verify". This is more difficult when verifying the decision process of an algorithm is unclear.