The legal rule that computers are presumed to be operating correctly (2022) [pdf]
benthamsgaze.org
benthamsgaze.org
Defense attorneys in a DUI case got their hands on the source code for the breathalyzer. It turned out to have terrible programming, e.g. calculating new averages by averaging a new value with the previous average. The case went all the way to the New Jersey Supreme Court, which still found the device to be acceptable.
With software that bad, who can say
the issue is that taken as a whole the quantization of the samples into 8 bins is a much bigger issue along with the problem of the no hardware watchdog or hardware malfunction alarms. along with the poor testing methodologies too.
If you have a set of readings, say, [0.1, 0.02, 0.3, 0.05, 0.08], normally when you average them you would get 0.55—the mean of the set.
Calculating the average by "averaging the new reading with the previous average" would mean new + old / 2 every time. That means that for each reading after the first, your "averages" would be: [0.06, 0.18, 0.115, 0.195].
If we add a new reading of 0.01 to each of these, in the first case, we would get an average of 0.46, and in the second case, 0.1025. As you can see, even taking into account the already-very-skewed numbers, the second case biases it much further in favor of the new reading (which, in this case, is very low compared to the existing readings).
1-Yes
2-No
3-Unspecified
Of course the average was around 2.011.
Hard to get upset over that. What matters is not the signal processing but the validation. You take a bunch of people with various blood alcohol levels measured by some already accepted lab technique and you verify that your new measurement technique is measuring within some acceptable error bound of that.
Which is not appropriate here...
I always knew about the theoretical cosmic ray bit flips. Before listening to this episode, I did not stop to think how often they actually cause problems.
[0] - https://podcastaddict.com/wissenschaft-auf-die-ohren/episode...
My methodology was simple. On a Linux home server that had plenty of spare memory (non ECC RAM) I ran a process that simply alloced a large buffer and filled it with a pattern. It would then periodically scan through the buffer looking for changes to the pattern.
I ran this for over a year which should have been long enough given the amount of RAM I was using and the rates that I found in the literature for cosmic ray induced bit flips resulted in several flips.
My method would have missed a flip if the page it happened on had been paged out sometime in the past and that paged out copy still existed, and the bit flip happens between the time of the last scan and the time the kernel decides to discard that page. On the next scan it would page fault and load the good page.
But the system was very lightly loaded and almost never actually had to page things out, so most of the time if a bit flipped it should have still been there by the next scan and so I don't think this explains why I saw no flips.
A few years later I got a Mac Pro, which had ECC memory. I used that mostly at work from 2008-2017. I got another Mac Pro in 2009 which I used at home from 2009-2017. I'd occasionally look at the memory status in System Report which should say "ECC Errors" if the ECC had to fix any errors and only ever saw "OK". I'm not sure if that resets on boot and I only looked occasionally so if it does reset than it is quite likely I would have missed an error statuses.
If you ran that machine in a hotbox at 85°C for a year I think you probably would have experienced higher error rates. Bonus points if you also force the data to swap out a lot so it gets transferred through as many data paths as possible.
Also, there was an expectation that comsic ray induced bit errors would grow as ram circuit features shrunk, but it ended up not happening; reasons unknown or at least I never saw anything suggesting a reason.
Getting errors in a small sample of RAM is unlikely, unless you specifically induce them by using debug features or misconfiguring your system (but some systems conspire to misconfigure themselves, making it easier to observe! a couple years ago, retail motherboards really liked setting the ram voltage too low)
It's not 3.6 roentgen...
Joking aside that's incredibly fascinating, I never thought that ECC memory has that much of a performance impact. Might be more optimal to just get a large jerry can of water and put that over the server as radiation shielding lol.
I had a similar issue one time where a Pentium III era server rebooted and came up with a comically small amount of memory, maybe 2-4 mb instead of 128 mb. That wasn't too bad, because it was a very lightly service; important to be on its own machine for reasons, but didn't need much. Just ran a little slow when it was running from swap. I think it did trigger a swap usage alert, and then it was like why is it swapping, why is it so slow, wait why doesn't it have any memory!?
I'm told it's the equivalent of a chest x-ray...
I see you, brother! :D
Jokes aside, that it seems really unlikely that cosmic rays would be clustered past, like, a couple hours, right? It isn’t like some neutron star is, like, tracking your server as the world spins (well, I hope not, I mean who’d you piss off for that to happen?).
Anyway, this fits my totally unscientific expectation that cosmic rays are just sort of like an informal description of hardware bugs that nobody can reasonably find beforehand. A server that seems to be hit by lots of cosmic rays probably has a dodgy connection somewhere inside it, but I mean maybe it’s the RAM, swap that out or replace it… but maybe inside the chip SOC, so what are we going to do, bust out the electron microscope to check all those connections?
In terms of diagnostics, we pretty much just asked for ram replacement, if that didn't work, cpu replacement, if that didn't work, motherboard replacement. If that didn't work, the chassis / rack position is clearly cursed, don't give us anything there again, please. :D I don't know what they did with the hardware we didn't like, maybe send it to the manufacturer, maybe give it to customers they don't like, maybe surplus it.
So, at scale, they were getting failures constantly. (It didn't help that the QIC tape backup was basically write-only.)
As a rule of thumb, you get 2 muons through your head a minute - but of course your head has a huge volume compared to memory chips.
In the case of a scientific simulation, we can find algorithmic speedups that are "lossy"- an example would be approximations to n-body systems, where you need to calculate n-squared interactions (between all pairs). You can calculate all n-squared interactions which produces teh completely correct result. But since atoms that are far away don't interact strongly (falls off as 1 over r squared or more). So you can maintain a neighborlist- all atoms within a distance R- cheaper than you can calculate n-squared interactions, with some tiny error that is unknown. It's assumed in many cases the errror is neglible and the speedup is huge.
Or you can switch to using a particle-mesh method which involves taking a fourier transform, doing some work in fourier space, then an inverse transform. The results are nlogn, and the error is small (and known). Speedup is signfiicant but takes much, much more computational skill and infrastructure.
Next, the case I'm referring to of an accelerator running a tensorflow job, it's a totally different scenario- here, some random subset of machines will repeatedly return the wrong result- say, for a matrix-vector operation. Maybe garbage numbers, maybe all zero, maybe some infs. When that gets summed into your gradient, it often causes blow-up and the entire job terminates. It's not clear whether it makes sense to make high-performance jobs have to be robust- I see them as special cases where you're working hard to make sure the computing substrate is effectively 100% reliable (by sending/fixing those machines).
Other folks have observed that some amount of small noise injection to the gradient can help training, but the sorts of errors I've seen almost immediately terminate the training job. I don't mind intentional noise addition, but noise due to hardware that is provably, reliably, and repeatedly miscalculating results? Not so much.
The best was a one-off error log along the lines of "unknown type System.DateTime". Huh? That's a system defined type that just went missing. Never saw it again.
Another at a different employer was a crash that occurred after a check condition that absolutely should have gated the crash from being reached. Single threaded. Simple microcontroller. Had to reflash it to flip the bit back. After doing the math on how much RAM we had in the wild vs. cosmic bit flip rates reported in super computers, we had to expect one flip per year.
If it's a safety critical system, server or not, use ECC RAM!!
Any reasonably competent software developer or engineer should know that just because the customer dashboard is green and everything is working as expected doesn't mean that there isn't an absolute dumpster fire raging in the background (at least, having worked in tech for 20 years across many different verticals, that has been my experience).
Not even covering the "kick it down the road" issues that sometimes (and sometimes don't) evolve into "features that aren't bugs" that have workarounds on workarounds on workarounds that will eventually have their own bugs. The stuff we know we shouldn't do / that isn't good, but that happens anyway.
I subscribe to the philosophy that computers cannot be held accountable, because they're only operating as intended, or at least as implemented. CentOS used to bill itself as a "bug for bug" compatible RHEL clone. I have always thought that to be a good perspective.
Humans are ultimately responsible for the results of computing. In my life that means I review the reports output by my ERP and sanity check numbers before I sign tax returns or remit quarterly tax payments. I trust that the numbers going in are correct, and I trust that (as was said up-thread) statistically my ERP is reliable, but at the end of the day someone has to be accountable for how the output is used.
He immediately remarked "they're scientists (physicists who sent the fax) and it was impossible that they wouldn't have accounted for that".
I've been a software engineer for 10 years. He's a well-read hard-working blue-collar guy, working as a taxi driver and behind a deli most of his career. I just nodded and moved past it.
People want to anthromorphize AI. People want to yield divine knowledge to computers. Any sufficiently advanced technology is indistinguishable from magic indeed.
I have a somewhat awful belief about humanity: that about 70% of people work jobs they’re not very good at. Most of the time this is ok, because so many jobs are bullshit anyway. Bad real estate agents still sell houses. Grumpy sales people still sell product. Social media managers - well, let’s not go there. But it has some strange consequences - like how most therapists are (in one study) less useful for the patient than journaling. Or in my opinion, the prevalence of crap software.
Ask about his work. Ask if his coworkers made a fax machine, if it would work reliably. If he gets it, you’ll know. You’ll see it in his eyes.
The latter problem is more important, but by lumping this together with "malfunction" and giving technologists basically a complete pass on the entire hard part, this kind of rule is a loophole wide enough to pass a jetliner through
I'm not a UK lawyer, but the law they quote says nothing about the logic machines are programmed to follow presumptively creating reliable evidence. It could be read to say that computers should be presumed to be executing the instructions they're given reliably, unless evidence shows otherwise. It's about malfunction, not misapplication.
Perhaps some of the Horizon case decisions showed judges improperly presuming that Horizon calculated correctly, and not just that the computers were running Horizon correctly. But the article doesn't show they did, or even explicitly say they did. Conflating two separable issues, it fails to address whether or why different presumption rules for each might be desirable.
" The Law Commission failed to address the strongest arguments against repeal without replacement, …. It ignored the advice of the experts they cited who all argued that the focus of courts should be on the reliability of computer evidence, …. The Law Commission’s comments and conclusions revealed that they had not understood the nature of computers and complex software systems as described in the sources upon which they relied.’ " https://www.computerweekly.com/opinion/The-cause-of-the-Post...
I mean, hey, the computer is operating correctly, it can't make a mistake, so give me my money and turn yourself in...
You are making a technical distinction the law does not make. The law relates to the output - to the evidence generated from the system as a whole. The presumption is that the logic is correct too - which is why it is such a terrible assumption.
In the specific case of the Post Office miscarriages of justice, the system as a whole included Horizon central office staff who made manual and sometimes incorrect adjustments to the individual Post Office ledgers, and the discrepancies were later blamed on the postmasters.
Extending this rule to LLM’s would clearly be disastrous. At a minimum, a record-keeping system needs to be written the old-fashioned way.
If the `alleged theft amount by defendant` == (unsigned int)((int)-1), that should cover those cases, at least in theory. The profs would just blindly vouch it, but at least it'll be a point.
edit: gender neutral language
The proposal is directly informed by the Post Office scandal in that it basically requires the prosecution to produce the same already-existing documents that revealed that Horizon had significant known faults.
Well into the 90s the general populace considered computers infallible, it really took mass adoption and experience with how frequently they crashed to change that perception.
The big win would in this case would be that when the vendor conspired to hide bugs from their own tracker, they would have been creating criminal liability for their employers. Which Fujitsu and their subs richly deserve.
Bonus: Many opaque systems would have to be aired in an open court room to ascertain whether their invariants do, in fact, survive scrutiny.
It’s also extremely easy for the computer owner or IT people to do the same.
But this is so very dumb.
It's not like they're going to give you the keys to the server room.
It's more than that. You don't have to instruct a jury that they have to accept that a person does what they're supposed to do unless proven otherwise... unless that person wrote code for a computer.
The alternative would be to say that only a system that has been proven correct could be used as evidence which would force pretty much all accounting back to paper.
Assuming that the paper accounting could be proven correct, when the computerized accounting could not, your proposed alternative would seem to be an upgrade.
You can make paper backups of course, but you can also make electronic ones, the electronic ones can be shipped around the world basically for free.
We're still a long way away from pulling that off with computers, but the idea that someone could have made a mistake shouldn't be that hard to grasp.
Don't pretend that fucking over innocent lives in the pursuit of profits is a law of nature.
"It is a matter of surprise that important documentary records, such as the Fujitsu Known Error Log (KEL), were disclosed only in response to a direction from the court and in the face of opposition by the Post Office." (https://journals.sas.ac.uk/deeslr/article/view/5240/5083)
The proposal in this paper would require prosecution to disclose extant documents related to the reliability of the system. In the Horizon case, such documents existed, but were never disclosed, as current practices did not require the prosecution to do so.
Are prosecutors in the UK required to disclose both incriminating and exculpatory evidence to the defendant?
But from other articles it sounds like these prosecutions involved some kind of parallel legal system just for the mail?
I think though that the issue that this saga hit is that the law was such that the computer “evidence” was presumptively correct so they tried to pretend any “bugs” they were aware of didn’t change that fact and so the existence of bugs was not “exculpatory” - obviously BS, but it seems like that was the core behaviour/belief of the post office. That put the victims in the position of having to disprove the accusation but the only “evidence” in the case was the presumptively true report from the buggy system.
Ie the only evidence presented was a system that was buggy, but the victim could not get evidence the system was broken without first proving it was broken.
Yay!
No. It'd actually just force authorship and disclosure of test cases.
No, there are many alternatives. Including, most critically, the one laid out in the paper under discussion. Which is – spoilers – not “only a system that has been proven correct could be used as evidence”.
The original rule-makers probably never considered the idea of a giant, bespoke enterprise software system; it's just a particularly terrible edge case that got caught up in what looks like a reasonable efficiency measure.
Eye witnesses, expert witnesses, etc. can be unreliable. It's very difficult to prove that they have not made any errors. However, we don't ask juries to presume they are infallible.
There is a difference between being infallible and being correct. No one assumes that computers cannot err, just that they have not erred unless there is reason to believe they have. Likewise, if a witness gives coherent evidence and no one has any reason to assume they are wrong or lying, that evidence will not generally be disregarded simply because humans are fallible and therefore the evidence is presumed to be flawed.
There's no reason to think that programmers are less fallible than other people.
> required the prosecution to prove that a computer was operating properly at the relevant time before a document produced by such a computer could be admitted as evidence.
But I'm sure you actually mean to agree, and that your comment was just garbled in transit. :)