How a software glitch at the UK Post Office ruined lives
cnn.com
cnn.com
If all people could be selfless, humble, intelligent, competent at everything they ever try to do, AND be people of integrity, I'm sure many of the world's problems would disappear, in addition to just IT problems. People suck, organizations suck, societies suck, and we all suck in our own way. I think organizational research like I'm trying to do still tries to point us towards a better path nevertheless, but sometimes it's a throw your hands in the air thing. Some people will never care.
Needless to say, I'm no longer focused on researching this topic, as it seems really well-researched already. But it's still interesting to see that this particular example still pops up in a news report now and then. There are still plenty of other big examples that pop up every year, but this one seems to have staying power to stay in the news media.
Just a few basic things that wasn't included, no audit/transaction logs, transactions modified by tech support to keep the system running.
Operators couldn't prove they didn't steal funds, and the british law that computers systems are to be trusted as fact, pretty much convicted them all.
There are two problems here. First, the branch manager is responsible for calculated shortfalls, even if the software is broken. Second, there is no way to overturn broken software. Third, the prosecutors are overzealous in trying to shut these people up and convict them straight away.
The software itself was just a convenient medium for abuse of authority.
What is "IS" in this context? I did some Operations Research modules at uni and thoroughly enjoyed it, but it had nothing to say about why projects didn't work.
It's enough work to do one's own job well already. The go getters can of course do it as a natural course of action, but they are outliers. There are a limited number of job opportunities that require developing this experience on the job, so there are limited opportunities to become a good boundary spanner in the first place. Furthermore, people aren't naturally interested in multiple disparate subjects. True renaissance folks like Leonardo da Vinci who are interested in becoming experts in both art and engineering are rare. Elon Musk types that will try to dive deep into multiple unrelated areas are rare. All of this adds up to boundary spanners being rare. As such, leaders who can develop cultures that handle multiple areas simultaneously (see the founders of Flexport who understand both tech and shipping logistics) are rare.
In short, expertise is hard to develop, expertise in multiple areas is rare, and coordination between two areas that understand different worlds is difficult without boundary spanners. As a result, you get failures. See any software engineer who creates a startup to try to revolutionize some old-school industry and then fails dramatically because they don't understand the problems that actually need to be solved. The outliers will figure it out, but not everyone can become an outlier due to reasons discussed, among other reasons.
The entire field of cryptography wouldn’t even exist if the boundary of research ended at “good actors”.
I am skeptical that there are software development practices that will allow me to hire a team of feckless incompetents and have them develop quality software. If you know of any I'm interested to hear about them.
But you are entitled to your opinion and that's fine. We can agree to disagree, nothing wrong with that.
That’s the point. Cryptography is significantly better precisely because the research effort has gone into systems that are hard for people to fuck up.
Just look at the fight was to get everyone to agree that the model should be, “everything including the algorithm should be public, except for the key”. That’s a socioeconomic argument.
> Even when the technical ideas are perfect, organizations still need to implement the ideas. That implementation tends to not follow allegedly perfect specifications for many reasons.
And that’s why making safe systems where mistakes are protected against is a critical area of research.
Rust is popular because it protects against whole classes of bugs, despite it being no faster than C/C++.
You can talk about best practices all you want. You're not going to get most organizations to afford or convince the best people capable of following best practices to come work for them. And even if they did, those best people will leave before they can even change the technical culture. No effective person would put up with the insanity that exists in subpar organizations, many of which continue to exist in spite of their incompetence for many other reasons.
You have no concept of working in the real world where people who suck exist. You talk like you've only ever worked with all-star Linus Torvalds types. Of course it's easy to do what you are recommending when you're working on the Linux kernel, for FAANG, startups with competent founders, etc. All those organizations are able to do what you recommend for reasons that many other organizations can't.
You're a Xoogler working with startups. I get it. You're in that world. You have no idea how to fix an organization like the British Postal Office so that they will do IT competently.
They didn’t have to do anything more than sign in to their account and they got better security than the military leadership of WW2.
It’s easy to lose focus and take “dumb users” as an undefeatable entity, but that’s the lazy way out. The most significant advancement of cryptography was asymmetric crypto that explicitly meant a moron leaking the public key didn’t mean shit.
You can't just drop an Instagram-like app into the British Postal Office and then everything's great. Instagram works as a standalone app that doesn't need to comply with any exogenous processes or standards. It can set the standard process for itself, and then all of its users need to adapt to it. There is no way to design an app outside of the British Postal Office that will fulfill their needs and then drop it inside of the British Postal Office and expect to work. Even if you forced the organization to reorganize itself in order to adapt to the app (which happens a lot), the problems will be inevitable.
Even Office 365 or Google Apps, stars in the SaaS space, require internal administration and customization when being used inside organizations and they can be misused. Something as simple as this person should be part of this security group but not part of that security group. Such misconfigurations are inevitable because organizations are messy.
People smarter than you and I have been trying to solve the problem of good IT governance for decades and have so far failed. And the problem has nothing to do with the quality of the software engineers who make the product, nor their technical decisions. It is orthogonal to the real issues. The fact that dumb users can use encryption today without realizing it has zero implications on solving the issues that organizations actually face. It has zero implications on how they use their technology. The only thing it's done is made the technology more trustworthy for transactions of information, but it did nothing to change work habits, decisions, or perceptions about technology. We know because there are studies on how people interact with technology.
These are not technological problems. They are human problems. Things as simple as "I am petty and don't like that employee" or "I'm gonna make sure that my friend gets to have sole responsibility for that app's strategic focus, even though he knows nothing about how to do that department's work, but he's my friend" or "I need this political win and that's more important than hiring the right technology experts or implementing the right feature the right way." They're simple problems to express but intractable to solve. They're intractable because they're emotional and irrational, spawned by people who need therapy. And most often, the people with these problems aren't stupid or dumb. They're actually often smart, which is why they're also often in the position to make the wrong decisions for the wrong reasons. And then it trickles down throughout the organizational culture.
Technology cannot solve this simply because technology can always be discarded or misconfigured, despite the technology's design. Nobody can force an organization to use a technology, especially when it doesn't have the necessary experts to implement it properly. The biggest problem was that nobody at the British Post Office cared for quality control of the system. Bugs continue to be found in cryptography. They're rare, but they are found. Then they are patched. A lot of organizations don't care and then they have security holes simply because nobody cares about patching. Such lack of care extends beyond just cryptography. Automated updated certificates like LetsEncrypt does not solve this problem because the problem runs deeper than keeping a certificate up to date or running automated security patching. Certainly, nobody can force an organization to use LetsEncrypt. Nobody can force an organization to keep the right people in the right security groups. Nobody can force an organization to disable network accounts for employees fired for embezzlement. Nobody can force an organization to care about documenting, reporting, and fixing bugs.
Being right on a technology level has no bearing on whether one can ensure that an organization makes the right decisions overall. Most of the most important decisions aren't even directly related to technology. Even if the easy solution is as simple as use a SaaS that is as simple as Instagram. The dysfunctional ones will say, "Screw that, I want my bonus or I want my job security or whatever, I'll make sure we never use that Instagram-like app, or anything like it." Or worse, they'll try to use it with the best of intentions and then still screw it up massively when they deploy it for employees.
If you can't accept that possibility, we have to agree to disagree.
I do, people are fucking hard to deal with an the entire field of cryptography is a technical solution to that. If you were at the helm it wouldn’t exist.
>There is no way to design an app outside of the British Postal Office that will fulfill their needs and then drop it inside of the British Postal Office and expect to work.
Yet it does every day and you just take it for granted. What voltage electricity do they use? Do they use phones? Do they use email? Do they use TLS? Do they use standard plumbing?
>These are not technological problems. They are human problems.
Human problems are for technology to solve. Keyword exchanges for ensuring Diffie Hellman secrecy are completely to deal with humans and the problem space it’s meant to address.
>Technology cannot solve this simply because technology can always be discarded or misconfigured, despite the technology's design.
Again, this is wrong. If the organization interacts with the public they are forced to interact with interfaces defined by browser manufacturers. It’s getting extremely difficult to run an insecure website that takes passwords without being flagged as such by systems outside of the organization’s control (e.g. chrome).
You can’t just refuse to implement smtp authentication, DKIM, and SPF anymore and expect to be able to email the rest of the world.
Your view is of the world as it existed during the 2000s. A giant buffet of technological choices, some good and some bad. That world is gone.
We now (for better or worse) have minimum bars set by a cabal of large internet companies if you intend to interact with the general public. Safari will not lower the bar for accessing the postal service, no matter how bristled the mustache of the politician who requests it.
The entire field of research you purport to not exist and not matter is what drives this cabal forward. They remove various footguns every day and take the reins out of incompetent operators hands further with every release.
Your argument seems to be “I’ve thought of a case where humans can screw up so therefore the entire field of taking the ability for them to screw up away is invalid”. It’s completely ridiculous and I’m pointing out that it is both active and making major differences.
If you don’t want to participate that’s fine. But telling yourself it’s because it’s an unsolvable problem is just a lie.
None of your solution philosophy would have prevented what the UK Postal Office experienced. Their issues had nothing to do with interacting with the public. Their use case was an internal black box where things went to hell and they refused to see that it went to hell because it looked like it was working.
You have no idea that we're talking about different things. You can just refuse to implement smtp authentication, DKIM, and SPF and not email the rest of the world. It is possible when you don't care about communicating with the rest of the world and you only use it for something it wasn't meant to do. And then weird phenomena emerge from that. Or you are emailing the rest of the world but you have no clue that you're having one-sided conversations because you don't care about replies. There are so many ways your assumptions can fall apart but it doesn't matter because they're not even trying to use the darn thing for e-mail.
You can design perfect technology in your perfect world but you can't force people to use it the way you're planning. In all your examples, you have the assumption that people will use technology for what you're designing it to do. Those assumptions mean nothing when they use your technology for something you didn't expect and it seems to work for them anyway. And there's no counterpart that they actually care about that tells them otherwise. You can try to explain to them that it's not working. You can even show evidence that their outputs aren't matching what they claim they want. But they won't listen because to them, it looks like it is working.
You need to stop bringing up examples and use cases that have nothing to do with the problems that the UK Postal Office was actually experiencing, and that all subpar organizations experience. Their problems are not the problems you are trying to solve in your logic. And even good organizations experience the same problems too, just to a lesser degree.
> The entire field of research you purport to not exist and not matter is what drives this cabal forward. They remove various footguns every day and take the reins out of incompetent operators hands further with every release.
I never said it didn't exist. I am saying it's not related. Again, the entire field of research that you champion solves a problem that is completely different from the problem that the UK Postal Office experienced. If you understand what happened with the UK Postal Office completely and then offer viable explanations of how your ideas would have prevented their problems and achieved their organizational goals, I will say I'm wrong. Hey, I'm not so arrogant to say it's impossible. But I will say that everything you've said so far is so unrelated to what their problems actually were at a root cause level. To say otherwise is a lie.
To be fair, I used to think like you. It was because I believed that we could create technological solutions to human problems that I didn't understand why organizations didn't just do technology properly. After diving into the research literature, I have realized that's naive. I no longer think like you. It is more complicated than what code alone can resolve. If you can't provide a solid analysis of how your ideas would have prevented the UK Post Office's problems, we really need to agree to disagree.
and yet you don't see very many buildings collapse, bridges fail, and damns flood down valley. The idea is that there should be liability assigned to important systems, for which this liability makes the onus on the creator/owner to build in safe guards, checks or other protections to prevent disasters.
Why this isn't applied to software engineering is a whole nother story, but i think it probably should. Move fast and break things is not something i wanted to hear tbh.
You're doing super well if you're successful enough that some of your values need to be re-written.
Computers, at its best existed for only 200 years. That's counting from the difference engine that Babbage never finished building, and if we are talking about the first modern computer that does contribute significantly in a way, that's around the time ENVAC and Z3 appeared, just around 80 years.
Structural Engineering existed for over 2000 years. With that comes with a lot of people died in natural disasters and advancing material science to improve structural integrity. But even the Japanese to this day still can't solve the earthquake and the problems the fallout makes.
> Why this isn't applied to software engineering is a whole nother story
I actually think it is. That's why real-time operating system and mission-critical hardware exists for human to fly in the air and explore the space, not counting a lot of time-sensitive industry robotics software that controls the actuators in real time as well.
That said, the strict requirements of real-time programming requires a lot of expertise such as CPU cycle counting, CPU slack reduction, timing requirements and choice of algorithm (such as EDF scheduling and real time memory allocators like TLSF, and you won't see jemalloc on embedded devices, right?), stack or heap memory allocation which also complicates the programming stuff, because removing malloc might free you from OOM but this means a lot of functions needs to accept extra parameters to state the output location, and that means a lot of normal software can't be used. (You can go for a hybrid approach by using arena allocation, but that still isn't a perfect solution)
As you see, even for soft real time engineering (which I actually mentioned so far, I don't know much about the real hardcore "hard real time engineering" though), the sheer complexity here already, means there's a lot of design decisions which simply just makes people stay away and just go for normal software engineering (but in the end, if everything is predicable its fine).
> Move fast and break things is not something i wanted to hear tbh.
When Zuckerberg said that, he's referring to his startup mindset when Facebook was really just a startup in his day. In a highly competitive market environment, startups has to fight desperately for their own survival, even with lots of fundings and VC rounds.
Startup can move very fast and their agility is the only weapon against the old dogs. Now you've become one of the old dogs, you don't break things.
Meanwhile, software's inner workings are encapsulated so that they're not self-evident as to consequences if things go awry. Furthermore, software is inconsistent in how it works, and this trains people to think that software is just naturally glitchy, but the glitches aren't a big deal. See the joke about different types of engineers also: https://www.reddit.com/r/Jokes/comments/pqr8t3/four_engineer...
The mechanical engineer says: “It’s a broken starter”
The electrical engineer says: “Dead battery”
The chemical engineer says: “Impurities in the gasoline”
The IT engineer says: “Hey guys, I have an idea, how about we all get out of the car and get back in”
If that's what the IT guy recommends, how do you get users, never mind corporate managers and executives, to take software system quality seriously? Obviously, this isn't what a proper software guy will think, the proper software guy knows that things are a bit more complex than that. But this is what the software guy communicates to non-technical people. And only a few rare non-technical people will tell the software guy, "OK, look, just tell me what's really going on here, how deep the problem is, and what we need to do to fix it, no matter what it takes." Most people won't have the time for that because fact of the matter is that the system satisfices needs until it doesn't, and then it's too late.
Software can, and should, also be helpfully smarter. Where there's a chance for invalid inputs, they should be validated. That would help with errors like:
* Authentication failed, check your account credentials.
* Could not connect to remote server.
* Invalid data input, failed to parse around byte/octet 1,234,666, near line 56789 (printable filtered) text: 'Invalid String Example Here'
Yes, those are my take on 'software' versions of the starter, battery, and impurity.
Lots of well written software will do this, mostly good compilers for software.
Poorly written software, often hides these errors if it even collects them at all. PHBs and other less professional people seem allergic to thinking, for them it's easier to call in a specialist or just not do it with the broken tool.
I have worked with ECU programming. When test drivers got stuck and called me it was the first thing I told them. "Restart the vehicle". If that didn't work, "disconnect the battery and wait 5 minutes".
They quickly learned.
2. Early implementations of Britain's TPWS on railways are exactly like this. Suppose you're pulling into a terminus station and, for whatever reason, you happen to stop such that your train's sensors are right on top of the TPWS "toast rack" (basically think radio transmitter). When the next driver turns the train on, it can see a TPWS transmission. Now, it knows perfectly well it isn't moving, so we're not in some terrifying near-death scenario, and it could let you, a human train driver, sort this out by, you know, not crashing the train into the station buffers. Nope, in v1.0 the firmware just considers starting in this situation to be a fatal error and won't let you move at all, the only authorised solution is to get rid of any passengers who've boarded your train, switch the train into it's (unsafe for passenger service) maintenance mode, drive it away from the beacons a few yards, stop, turn off the power, and then reboot the computer where it now can't see a troubling TPWS transmitter and now it's a working train again.
"The IT person says". IT is not an engineering discipline (nor is software development in most cases), although they really should be.
And "turn it off and on again" is the canonical "joke".
This sort of thing is inevitable once a company is controlled by people who don't understand what the company does.
Management replaced know-how with software, and it mostly worked. Then they probably fired the few remaining people that had the know-how. Then, when it became painfully clear something was wrong, management chose to blame anyone, including to the point of pursuing prison sentences, rather than admitting they didn't know their business and someone made a mistake. They chose to destroy people's lives rather than admit they couldn't do the ONE job they claim they can.
I think we should examine why this doesn't happen with structural engineering firms.
For the third time on this thread I'm going to say it is because you have to have proper qualifications to do that for a living.
And the enormous Millennium Tower in San Francisco is sinking and tilting.
So it's not unique to software.
At 10pm, a janitor came in, unplugged the computer, and plugged in his floor clearer, cleaned the floor, and then when done, unplugged the floor cleaner, and plugged the computer back in.
The consultant then suggested using the outlet on the other wall that had a free outlet (and got a 'difficult to unplug' cover for the outlet with the computer).
The next night he stayed late again and the janitor used the outlet on the other wall.
The consultant then told management that the problem had been solved and it was a buffer problem.
Apparently there are tons of these stories [2] [3], many of them probably urban legends.
But there was one last year that was definitely real [4], when a cleaner removed power to a freezer holding decades of samples, apparently because they were annoyed by an alarm sound it was making.
[1] https://www.goodreads.com/book/show/3227607-the-devouring-fu...
[2] https://old.reddit.com/r/talesfromtechsupport/comments/5yrs1...
[3] https://www.logikalsolutions.com/wordpress/information-techn...
[4] https://www.theguardian.com/us-news/2023/jun/27/cleaner-coll...
https://archive.org/details/devouringfungust00jenn/page/96/m... (pages 97 to 98) is the proper telling of the story.
There is legal accountability though, and a much better understanding of who can be held to blame for what. Software suppliers are pretty much immune to consequences of carelessness.
However, I think the real problem is that software is far more complex than a bridge or a building. IT system have complex hardware (far more complex than any mechanical device) running even more complex software (counting all the layers from OS up). On top of that people (meaning users/buyers) have no idea how to evaluate safety/reliability/security and mostly seem to regard it as a nice to have to be traded off against other nice to haves, not an essential baseline.
Once its implement everyone assumes the computer must be right, and acts accordingly. The presumption that computers are right is even enshrined in UK law (the intent was to stop people getting out of things like speeding tickets by claiming speed cameras where faulty) but I think everyone has come across situations where the final word on a dispute was "the computer says so".
It does not get done everywhere because it is not a high enough priority.
The English term "Engineer" can be used by anybody though.
As time goes on, I feel more and more strongly that we shouldn't compare software engineering with other types of engineering at all. My younger brother works in the structural engineering space and the types of issues he's faced across different roles is very different to the types of issues I face in software.
Building design and construction have to handle the physical world and dealing with the realities of physically building the structure. Once it's built, however, the occupants don't make drastic changes to the building itself. In comparison, something like Facebook or Twitter needed to change drastically just to remain usable as they grew. "Just stop adding users" isn't a sensible solution to the problem.
Just to be clear that I do not excuse the, frankly, shit software design of Horizon and the fucking appalling behaviour of the Post Office throughout this scandal. I do think, however, that comparing software engineering the other types of engineering does a disservice to all and doesn't take us closer to actually improving the craft.
The structural engineers take a piece of steel out into the desert and they dump sand on it until it collapses. Then they divide whatever weight the sand was by 3 and write it down into a book which is then published for all the other structural engineers. This type of steel with this type of configuration with this length can hold up to this much weight.
Meanwhile, in software engineering the limiting factor that prevents collapse is often "these sets of constraints and requirements are too complicated for the team to keep track of in their heads in this messy codebase". Not only do we not have a way to test this and publish the results, it's not even obvious what exactly you're even trying to measure.
[Cyclomatic complexity isn't very well correlated to defects (iirc lines of code is better correlated). Design patterns, best practices, and code smells are all just poetry that we use to story tell our ways out of blame. And Weyuker's 9 Properties are at best a waste of time. There is currently no way to measure bad code that fails under load until it does so in production.]
FWIW, techniques to perform such measures have existed for ~30 years in PL research (some of Rust's type system comes from this kind of research), but I don't know of any industrial code that makes use of it.
So far the only[] thing I've found was Weyuker's 9 Properties from measurement theory, and that did not seem particularly compelling.
Now, I get that linear/affine types and dependent types (and theorem provers, etc) can be used to prove that you're following the spec. But this doesn't prove that the code is easily comprehended by the people interacting with it. Code that is provably correct (even in the event where the spec is actually exactly what you want) can still fail because it's too complicated for anyone to modify it anymore.
For example, the function composition operator type in agda is kind of intimidating:
_∘_ : ∀ {a b c}
{A : Set a} {B : A → Set b} {C : {x : A} → B x → Set c} →
(∀ {x} (y : B x) → C y) → (g : (x : A) → B x) →
((x : A) → C (g x))
[] - There are also random blog posts et al scattered throughout the net, however I haven't found anything that seemed to actually work. For example, https://www.sonarsource.com/resources/cognitive-complexity/Weyuker and cyclomatic complexity seemed to be the only things from academia.
Can you mention the features of these languages that you think are worth looking into?
Granted, this doesn't fulfill all the requirements for being resilient, but if we manage to make it usable (let's say as usable as Rust), it would be a pretty large step forward.
These people have been called 'software developers' or just 'programmers' back in the day. In fact, I'd argue that most commercial software development is more like plumbing than engineering.
Here's a hint: if you're writing control software for airplanes, medical devices or industrial plants, you're an engineer; if you're developing a UI frontend for a website, you're probably not.
I've worked on enough projects with "real engineers" to see that rigor varies significantly with the person. e.g. I've seen a dropout with more engineering rigor than a Waterloo grad (granted, this was at a company with a very selective interview process...)
In practice, you do whatever and then pay a licensed PE with an absolutely massive liability insurance policy to stamp your design. If something goes wrong, they take the fall and go sip drinks on a beach somewhere.
As usual, incentives rule everything around us.
So the US didn't fall far from the tree. It's just like 'daddy'. <wipes away imperial tear>
They kinda have - check out the cladding scandal [1]. It's pretty major, a lot of properties are essentially unmortgageable until they get inspected to make sure they're safe.
[1] https://en.wikipedia.org/wiki/United_Kingdom_cladding_crisis
There is still hope for Grenfell prosecutions after the inquest, but justice repeatedly delayed for no good reason is justice denied, obviously.
You used to have to get your design cleared by a Local Authority Buildings Inspector. Now you can opt out and get a private company to do it!
Some were deregulation, but some were outright deliberate fraud, like when Kingspan and Celotex scammed the test process and produced misleading safety documents to get their flammable insulation used where it wasn't legal.
And local authorities don't come out of it well either - the building was owned and maintained by a local authority, they hired the architects, the design-and-build contractor, and the guy whose job it was to carry out fire risk assessments. And they chose a bidding process that drove the price so low the bidders who wanted to use non-flammable cladding all dropped out.
Definitely need to stricter regulation though - the idea we could trust the construction industry to use flammable cladding safely has proven false. Flammable cladding should be banned on tall buildings all together.
Again, how is it you are allowed to do your own testing, rather than submit samples regularly to a test house? Or perhaps products on higher risk buildings should be spot checked. A sample from the jobsite itself could be tested.
Self- regulation is a joke.
Just you know, it turns out while the test rigs were being built, the BRE might have looked the other way while someone put fireproof magnesium oxide boards over the temperature sensors.
And when the manufacturer sent 'all the test results' to LABC for the certification, they didn't include the tests that ended in "a raging inferno". You see, that test was terminated early, so it wasn't a completed test. And when the manufacturer got the certificate granted, which included a bunch of caveats about how the material specifically had to be used, they immediately quoted it in their marketing materials without any of those caveats.
Of course all this fraud was done by some 18 year old junior employee, who was told it was 'industry standard behaviour' by a mid-level employee (who has conveniently since died) and the CEO had 'absolutely no idea' this was going on right under his nose.
Cant stop laughing. For some strange reason this sentence reminded me of Yes Minister.
Show me a similar story about stuff this bad happening at Google and I'll believe that there's something wrong with the field as practiced at its best.
This is why the CEO was rewarded with a CBE and a cushy NHS post in 2019 when the scandal was in full swing. She was acting like a good girl for her political masters in managing the crisis, stringing it out long enough in the hope that everyone would forget
the problem is people asking idiots who cant program even a simple button, to program serious enterprise, government, and military systems.
Hi! Let me tell you about my wonderful country. Norway! I'm sure you've heard about it. The socialist utopia in Scandinavia! ;-)
The Braskereidfoss dam failed during a flood last fall. Luckily just a 'river dam' and not a reservoir dam. Still rather bad.
Oh, and during the same freak whether event, but different place in the country. The Randklev railway bridge failed over the river Lågen, just by Ringebu.
And during the last 8 years, we've had two road bridges collapse. One by Sjoa in 2016 (Perkolo bridge) and then 16 months ago Tretten bridge. We had the fantastic idea of building wooden bridge sfor road traffic. Here's a nice article on Tretten bridge: https://en.wikipedia.org/wiki/Tretten_Bridge
But yeah, not too many building collapses . Just a bunch of infrastructure collapse.
When software breaks you restart and it's back to "working" order. It would be a different scenario if the computer would set itself on fire on failure.
Moving safely means moving slow, and that costs more - it's as simple as that. Plus it would also automatically disqualify like 80% of the workforce, due to the need for a proper formal education and certification, which would literally paralyze the industry.
And the current solution is not that different from what we have e.g. in civil engineering, the building codes for high-rise buildings are way more strict and complex than the regulations for a garden shed - in many places no permit is needed at all for it. And sometimes that means that that shad will fall in a catastrophic way, too.
Would that be so bad, in the grand scheme of things?
But you're right, it generally requires you have a 4-year degree and have worked under a licensed engineer for some number of years.
Them hiring only people with degrees is almost certainly orthogonal to legal requirements in such cases.
Any sort of professional interaction with patients, sure, whole different ballgame.
Components at every stage of the design and manufacturing of a commercial airplane are signed off by licensed engineers who are legally liable for negligence, and yet Boeing could outsource the 737 Max's software to $9/hr programmers in India. Obviously physical engineering failures still happen, but it's ridiculous that there's no liable professional responsible for software that keeps a plane full of hundreds of people from falling out of the sky.
If a licensed accountant was caught lying about hundreds of people stealing from their employer to the point of getting some sent to jail he could easily do jail time for negligence or perjury. Instead some seemingly anonymous dev wrote shoddy software that ruined hundreds of lives and the Post Office gets to just say that the computer did an oopsie.
Because bits don‘t rot and while you need to build a new bridge now and then which will include your learnings from the old one, you never have to replace software because it can be copied ad infinitum.
It feels like there is no ground truth in computing, each layer has to assume the layer below is generating random errors and try to account for that, because humans are not perfect.
The laws of physics stay the same, technology is constantly changing.
Because those engineers are required to have rigorous qualifications. Please see my comment linked below
Once a physical thing is built or being manufactured, the cost of change is exponentially higher than with software.
So software engineers can experiment more frequently and can get away with low quality, as it can always be fixed later.
Of course, this same mentality can carry into industries where the stakes are life and death, which is where accountability is very much lacking.
It’s that way because software is much much much more complex than that.
Building is a good example. One architect after university has enough knowledge and mental capability to design a building that won’t topple: There are discrete number of parameters to work on and fallible elements are identified and understood very well. There’s manageable amount of degrees of freedom that architect is working with. It’s also not very common for physics to get an annual update.
But when you go to funny-cat-pictures.com there are layers upon layers of complexity. There’s TCP, UDP, traffic encryption, indefinite number of hoops in routing, indefinite number of electronic devices with various quality, etc. Request comes happens from millions lines of code operating system and generation happens on a different millions lines of code of operating system with thousands of lines of codes of software code that interacts with databases, proxies, CDNs, load balancers and is templating, serving, translating, transpiling, compiling and serving while defending against adverse attacks, managing the cache, optimizing for network topology and requester rendering software. That changes every single day slightly.
In 99% the worst case is - your picture won’t load. But when you get to the more serious software development (in a way that it’s critical domain, not that “serve cat pictures” isn’t serious job, mind you) all of it is very very visible.
Every element is so unimaginably complex that it has literal tomes of knowledge written and published about it. Describing every element of transaction between end user and cat serving website would take hundreds of pages - and it would be different than other cat website.
Most of the building, bridges etc. was made either by one person or small group of people working on it.
Rarely software is written seldom, and many well known products have tens of thousands of inconceivable smart engineers and we still can make fun of their - not so rare - failures.
So yeah. Systems suck, organizations suck, people suck. But they suck in relatively safe but also complex environment.
For years software engineering was made fun of because it’s not real engineering. But when you look close enough the environment is orders of magnitude more difficult. Civil engineers might be offended but they can’t hold a candle to structures software is keeping straight.
Software is a bridge built on a raft, floating on an ocean, which tries to get shot down by armed pirates and semi-controlled by a cost-saving manager trying to look good in annual progress report.
> A: Because I am very, very smrt and engineering is for babies.
Lol. Sounds like someone was very upset about being called "not a real engineer".
Maybe that's survivor bias? Crappy organizations fail at infrastructure projects before they even get off the ground.
Conversly softer endeavors involving just money, people and services can operate for a long time on finger crossing.
Money, for one thing. Actually building the thing IRL is the expensive bit, but it's effectively free for software. So engineers are incentivised to get things right the first time, and software developers are not. It's much cheaper to skimp on QA, let users find your fuck-ups, and push an update.
This case was brought to public attention and repairs were attempted only because it's huge and involved hundreds or thousands of people.
How many disparate cases there are, where people's lives are destroyed and innocents are rotting in jails, we have to ask?
Judges, salesman, and managers don't understand that.
"We need some software, ok let's get a big reputable company in to do it for us, we shouldn't get bogged down with all those horrible technical details"
They do understand that. They do not care.
The skillfull programmer may accept liability when you give him a verification team with a few PhDs, the ability to withhold signoffs, flexible deadlines etc. etc. Few are willing or required to pay for that. So they get a mystery box with a 90% chance of crap.
I think for the courts the issue is a bit more subtle. The question is, who's job is it to prove that the other person is wrong ("burden of proof")? Should it be the job of the prosecutor to prove that Intel's processor produces the right answer when an ADD instruction is executed? Or should it be the job of the defendant to show that Intel's processor doesn't produce the right answer? What about proving that the compiler produced binaries which faithfully represent the algorithm? What about Excel?
In our normal life, if a computer is doing the wrong thing, we don't start by assuming a broken compiler; we start by assuming that the new, not-well-tested code is probably broken.
It seems that in the UK before the 90's, the burden of proof was always on the prosecutor to prove almost everything about the system, which is kind of ridiculous. So they passed a law trying to fix it, but messed it up the other way, putting the entire burden of proof on the defendant, without giving them any real way to disprove it. (I mean, shouldn't "discovery" at least mean I can inspect the source code?)
A more balanced law would say that widely-used software with extensive test suites can generally be assumed to be working properly; but that custom-purpose software needs at least some level of evidence that it's correct, and that defendants have a right to inspect any software that's used against them in court for defects.
Laypeople generally understand that software may crap the bed in the sense of "the system is down, please wait, then try again". But few people have experienced subtle changes in stored data.
A judge looking into his document cloud may be ready to see a "sorry, not available right now" notice, but doesn't expect that some sinister program is, in the background, silently editing texts of his judgments and pronouncing people guilty when he intended to free them etc.
The problem with the Horizon scandal is in this sinister manipulation of data. It may also have been done by Fujitsu people themselves, in order to cover some tracks and tamper with evidence. This is a very untypical failure mode.
The legal system failed these people horribly. And the people who pursued these cases with no direct evidence whatsoever should suffer jail time.
However, the role of software here must not be minimized. Software makes it easier than ever to diffuse responsibility and create opaque processes that leave the least powerful people at the bottom of the hierarchy holding the bag. By rigidly encoding flawed assumptions and executing them without question, software is the ultimate realizer of our Kafkaesque nightmares.
No argument that the software bears fault, too.
However, when you accuse someone of stealing money, you should have to prove that they stole the money. This isn't some invisible crime. There should need to be evidence that the stolen money went into their account, got spent to buy something, got pulled from the till on camera, got transferred to Bitcoin--something.
The fact that all these people got convicted with no evidence that the money was ever in their possession is a gigantic legal problem.
I disagree, the software glitch was the problem here.
We are supposed to be able to rely on computers to store and add numbers or report a system failure. This accounting software showed in black and white that some funds that the sub-postmasters were responsible for had gone missing.
What else was the legal system supposed to do? The broken software was simulating crime perfectly.
Except it wasn't; the main problem was how the PO was handling it. ICL/Fujitsu were aware of near-identical bugs in an earlier project[1], and PO employees omitted parts of an audit from 2004 that described similar issues as well[2]
It all goes back to ICL/Fujitsu and the PO being aware of the issue and withholding the information from anyone not already "in the know"; lawyers, judges, changing witness statements to hide incriminating evidence, etc.
Our current system of the world quite strongly disincentivises honesty and integrity - rather, being a bombastic charlatan with a flexible relationship with the truth will get you anywhere.
In the early 2000s there was a TV ad campaign for “The People’s Post Office” where the sub-postmaster role was played by John Henshaw, a character actor known for playing hard bastards and, in his most recent role on The Cops, an exploitative bent copper from Bradford. A strange but apt piece of casting.
Imagine you are talking everyday to weird, bitter, arrogant, rude customers for years - even with high salary you will be not so positive.
So everything checks out.
Because we’re not professionals. We don’t profess anything and do not have standards. There is no regulation for our industry and no IT association that can strike you off from practicing this craft. There is no accountability, and when there is no accountability, people naturally regress to either lazy or exciting behaviours.
Our relatively small skunkworks team developed apps that changed the end-to-end solution delivery processes for major business units, both consumer and business sectors, saved the company 8 digits in opex and capex each year, and won an international award for "Best Support Team" (the Stevies, sort of known as the Oscars of the business world). Our greatest feat that year that enabled us to win the award was keeping the company afloat during a four-month union labour dispute by improvising solutions that automated everything in sight. At the end of the labour dispute, the CEO send a company-wide email about how important we were, awarded us this made-up award "Holding the Fort". When the union came back to work, we trained them how to use the new tools, but we unfortunately were also enablers of heavy downsizing, which I always disliked. Some of these people were hardworking people who did nothing wrong and followed the rules. Many of them were elderly and had little chance to go back to school to get new skills (we're talking 50-year-old clerical workers, etc). It drastically changed how I thought about corporate software work. That being said, we were all young cowboys, and it was possibly the best team I've ever experienced in my life.
I experienced the absolute opposite in many ways when I worked overseas for IBM, managing projects that spanned the Asia Pacific. I was the go-to PM many of their mission-critical infrastructure projects, including helping with datacenter migration from Japan to Australia, necessitated by the 2011 Fukushima earthquake and tsunami. I also experienced a middle ground as a venue technology manager for the Vancouver 2010 Olympics. Lots of pressure and set processes, but a lot of extremely competent people too.
Look, I'm not doubting your experience, but I've had mine too, which shaped my views, just as I'm sure that your experiences have shaped yours. We can agree to disagree, nothing wrong with that.
I can think of two things that I believe would make a difference in any LargeCorp: First, a standarized way to visualise and execute business logic that allows developers and management to reason together. (The no-code movement is on the right track in fostering a common way to interface with code). And second, a responsible editor for each piece of code.
I think a key factor is that software historically hasn't enjoyed industrialisation to the degree of hardware (or construction for that matter). I can buy a standardized CPU of millions of transistors and integrate it into a standardized motherboard with just a snap. We have managed to standardize software up to the OS level, but after that it's up to the developer and her shortcomings.
https://www.codevalley.com/ does some interesting work.
You will find that people just suck no matter what field you work in.
>Of course, I quickly found out that IS research had already figured most of this out, and that perhaps, people were just people and crappy organizations were just crappy organizations, and perhaps that's something that will never change because bell curve distributions exist for almost everything.
Hence why we need to keep things simple. The human part will never change, or at least change at rate that will take many generations to improve if you are an optimist. I actually prefer things to be Hybrid rather than all-in digital.
"One member of the development team, David McDonnell, who had worked on the Epos system side of the project, told the inquiry that “of eight [people] in the development team, two were very good, another two were mediocre but we could work with them, and then there were probably three or four who just weren’t up to it and weren’t capable of producing professional code”."
(Just in case somebody says I am putting blame on developers) Obviously, the responsibility is firmly on management. People making code bugs should not be held responsible for other people going to prison for it.
> In fact, staff at Fujitsu, which made and operated the Horizon system, were capable of remotely accessing branch accounts, and had “unrestricted and unaudited” access to those systems, the inquiry heard.
This has always bothered me. Sure, it's possible to build APIs that audit access completely. But I can easily write code that circumvents those APIs. Code isn't like a building where the walls are impenetrable and the doors the only possible access points - we can redecorate without ever touching the door. Building in an unaudited backdoor for operators seems bad, but if you can edit the source code the backdoors are infinite.
They were not able to do the first thing about running a transaction (ensure that one side of the transaction isn't executed multiple times). What you are saying is an obvious thing and yet it probably is well beyond the maturity of the team that was working on it.
They were using dial up ISDN lines to send the data back, but Riposte didn't support that, or scale to 20k terminals, so that was all new code
In general they had a distributed database that couldn't do ACID
https://www.postofficetrial.com/2019/12/fisking-horizon-tria... https://www.computerweekly.com/news/252496560/Fujitsu-bosses... https://www.benthamsgaze.org/2021/07/15/what-went-wrong-with...
This is a controversial opinion but I disagree, at least to a point. Managers don’t really know what we do. The only people who really understand the engineering trade offs involved are engineers. When lives are on the line as a result of our work, we shouldn’t be insulated from the consequences of our choices. That’s not good for society and ultimately not good for us. We change the world with our work. It’s healthy to understand and own the consequences of that.
The law agrees in parts. The principle of tort law is that everyone is responsible for foreseeable harm caused to your “neighbours”. Your degree of responsibility - and in turn liability - scales with how much expertise you have in the domain. An expert should have been able to foresee the harm more than a novice. The senior engineers on the team should have done better. I believe they are at fault.
(IANAL, this is not legal advice, yadda yadda)
That looks a lot like incompetent management.
My point is that particularly in the UK we have this culture that the Geeks should just do their job and let us Business Types take care of the rest. Countries like Germany have a much higher respect of technical people and qualifications e.g. it's very common for CEO's to have PhDs
In software, it’s the same. Your employer hires you because they need an expert. It’s your job to take responsibility for the software you write. Even if the CEO has a PhD, it’s not the job of senior management to review your code. And - trust me - they don’t want to. Instead take responsibility for your own work. You are worth your paycheck because you know how to build software well. Stop trying to shirk your job.
Most companies nowadays are more like restaurant owners i.e. computers are a core part of their business. They can't and shouldn't rely on engineers to get things right all the time because humans are flawed and therefore their code will be. The CEO has the responsibility to put people and systems in place to account for this.
Of course, if the CEO defunds their infosec teams, they bear the blame when they get data breaches. But I also don't think the engineers on the ground get to avoid blame when their crappy software leaks customer data. "I thought about making it secure but we had a deadline so I didn't" is a lazy excuse. Blaming management for everything is a lazy excuse. We're engineers. Not typists.
To continue the restaurant analogy, if customers are consistently getting food poisoning due to the sloppy hygiene of the kitchens then the buck stops with the restaurant owner. They'd need to diagnose the issue, put in place better training and supervision of the cooks and have systems to regularly check that standards are being met
The title of a singular "glitch" is very misleading in this case, as this was a cacophony of cock-ups and cover-ups for 20+ years
Also, a good manager should detect if a team contains a few "dead weights". It eventually always lead to frustration from more competent team members and this is something that can be detected and addressed if you talk to your team members.
The only thing that may be difficult for managers in the public sector is firing people unless they are contractors. I don't know how it works in the UK but I have worked in gov agencies in another european country and firing someone who was just working badly (to the point of wasting time and energy of others) was almost impossible. And sometimes upper management would throw unfit people in your team just because they needed to put them somewhere. I had a few time wasters in some of my team that had been put in our team by the unemployment agency. Basically they had spent money on them taking "classes" (which were just Microsoft certification classes) and threw them at us while most of them weren't even interested to begin with. You really had to do direct fault to be fired and it would take months or even years. The only one I have seen being fired was a project manager who booked fake meetings to go play golf during office hours.
A manager that has no technical knowledge is useless, like a jockey who doesn't see a difference between a poney and a horse or distinguish a healthy horse from an injured one.
Thats taking it way too far. Sure; its useful for managers to understand programming concepts (deployment, testing, etc). But the job of a good manager isn't managing code. Its managing people. Making sure Sally is happy in her new role. Setting up a meeting between Jake and the sales team so they can get to the bottom of that important bug. Helping mediate that conflict between the software team and the design team over what features to prioritise.
A manager is hired to be an expert at humans. Not an expert at computers.
Well, the reality is usually at these large software development agencies that senior engineers are prevented from doing what they think is right.
For example, they might have been pressed to deliver new features in an extremely inefficient system. They might have been inundated by low quality code from less experienced devs. They might have been busy with communication with stakeholders and unable to do much about it.
So singling out these developers would be like singling out couple cops for being racist when the entire Police department is known for racism. You know, technically you are right but that still does not seem to be right thing to do in case of a systemic problem.
The question would be if these developers had any way of knowing the consequences of what they were doing.
Also, "good developers" is a relative term. In an organisation like that a "good developer" might simply describe a person that is at all capable of writing working code. It does not mean they were experienced or aware what is happening around them.
It is the responsibility of managers to recognise the issues with the environment their people are working in.
Can you imagine how that conversation would go if the engineers were personally liable? “Hah you want me to sign off on that? No - I don’t want to get sued when it inevitably goes wrong. I’m sorry boss but I won’t do it. And you won’t find another engineer in the building who will. Lives are on the line. We either do it properly or we leave it alone. I’m not sticking my neck out to make the business a quick buck.”
> Also, "good developers" is a relative term.
Legally, as I understand it the courts look at job titles, education and experience to make a judgement.
If I’m going to be legally responsible for software bugs I must have the legal right to tell management that their timelines are not possible, and that they can’t deploy software I won’t sign off on.
That and for outsourced software the executives become personally responsible.
That strategy only works if the software itself is inconsequential.
Who actually works in a job like that? Do you seriously think you’re a slave to your boss, with no personal agency? Do you think your employer wants you to be feckless? Do you think that’s good for your career?
Your capacity to take responsibility is the differentiating factor between junior and senior engineers. Learn to step up. If nothing else, your pay check in 10 years time will thank you.
You position is a position of principle. When peopke lives are at risk (boeing, fukushima and yes even the post-office) pragmatism must prevail. Those in charge must pay the price, otherwise you incentivise financial results over everything else. People die? Oh that's because Greg in engineering is an idiot, burn him!
So if a cop gets an adresse wrong and is a bit too trigger happy and ends up killing innoscent people. Its their chief of police who should go to jail because they told them to go arrest a suspect? Unless we change the system to allow cops to just do whatever they want whenever with not leadership?
The idea that just because a programmer doesn’t have complete autonomy over their work that they suddenly become unaccountable for negligence and errors is ridicules.
Uh.....
> The only people who really understand the engineering trade offs involved are engineers.
Maybe you should have engineering managers who understand the subject?
Why not? If I'm the single developer and seller of this app, should I not be held responsible? What if there's also a QA person? Two of each?
Should the person selling or marketing the app be held responsible instead, even if they aren't technical? Why are the developers who didn't care enough to double-check their code free of responsibility?
In that case you are also the responsible manager or product owner.
> Should the person selling or marketing the app be held responsible instead, even if they aren't technical?
Of course. The person who takes the customers money is responsible for delivering the result and any warranty.
> Why are the developers who didn't care enough to double-check their code free of responsibility?
They are not the product owners. They don't decide what is the correct way the product works. Maybe they created the bugs because they implemented the specification exactly as written?
Of course we are, we're the ones writing it.
> Maybe they created the bugs because they implemented the specification exactly as written?
If your argument is "maybe they were told to write the bug in", I don't know what to tell you. If I were told to write a life-destroying bug into the software I worked on, I'd quit, because I don't want that on my conscience.
Like, you as a single seller are also responsible for making false claims. But, the justice itself should be more robust then that.
See the other discussion on HN front page about coding tests in interviewing.
Horizon was a child of the 90s. The software industry has changed a lot since then. Back in those days only Microsoft was routinely requiring programmers to code during the interview, so hiring was nearly random. Software teams often looked like that, with a tiny number of people who could write working code covering for many more who just couldn't at all.
The Daily WTF dates from this time. It's full of stories like the Brillant Paula Bean:
https://thedailywtf.com/articles/The_Brillant_Paula_Bean
You don't hear stories like that much anymore. The industry settled on testing concrete skills before hiring, and that washed out a lot of the people who previously managed to get hired into projects despite not being able to code properly. Whether Fujitsu does it or not now, no clue. But that situation wasn't unusual back then.
I don't think that was so much the case in the 90's, attitudes seemed to have been that these sorts of systems don't make mistakes, and therefore can always be trusted. I look at this as the 90's version of "The Titanic is an unsinkable ship"
EDIT: I would add that you can rely on systems / single data points for some decisions. How I see it is trust in a single point of data goes down as the criticality of the decision goes up. If someone sends you a message to meet them for lunch downstairs, that decision has low criticality and therefor you can rely on a single data point / application to make the decision. However for more critical decisions (ie. should we send this person to jail for stealing money), trust in any system should be low by default and a consensus from multiple data sources is required.
Um, I'm working as a professional computer programmer today, in 2024, and I can assure that programming continues to be replete with incompetent programmers who are not capable of producing professional code.
The IT bug was an issue, sure, but the political mismanagement of an institution stuck in the past is what caused all the ruin for so many people. And it flew under the radar until Netflix made a movie. Actually, the lady running the PO was awarded a Goverment recognition.
IT and code generation is full of pitfalls, but this one lays somewhere else.
__The state-owned Post Office acted as investigator and prosecutor in the cases, using the general right in English law for any individuals and organisations to pursue private prosecutions without involving the CPS.
A public inquiry into the scandal has heard that the Post Office, among other aggressive legal tactics, accused sub-postmasters of theft to pressure them into pleading guilty to lesser charges.
The CPS has identified 11 cases it brought against sub-postmasters that involved “notable evidence” from the Horizon system.
Legal experts said the government had been warned several years ago that private prosecutions carried a higher risk because those pursuing them were more likely to have motivations other than securing justice.
Lord Ken Macdonald KC, a former director of public prosecutions, said: “If you’ve got a body with skin in the game [such as the Post Office] acting as a prosecutor, that creates obvious risks and dangers.”__
Power corrupts; absolute power corrupts absolutely.
You were misinformed.
> As for the argument, "nothing else done otherwise", we don't see similar prosecution privileges for other essential services in life (health service, public transport etc.).
Yes we do; as the article says, in the UK any private organisation or individual has the right to bring a prosecution.
As I could be now. My sources were the BBC and PrivateEye podcasts. What are yours?
The CEO managed the crisis, therefore she was rewarded. The government should have shown leadership and demanded answers after the first few years of warning signs. Yet it took 20 years and a TV drama to force them to show any leadership
The UK Goverment really needs to shake off all this archaic nonsense and modernize themlselves. It’s not in line with the fantastic culture of their country.
No doubt pushed forwards by the racists believing in their heart that some of these sub postmasters HAD to be guilty because of their ethnicity.
I agree it's pretty stupid, but it is becoming the case for more and more people; while slightly different than the situation in the UK, forced arbitration clauses strip the right of the employee to seek justice, and they're getting more and more common.
> The Post Office threatened and lied to the BBC in a failed effort to suppress key evidence that helped clear postmasters in the Horizon scandal.
> The Post Office's false claims did not stop the programme, but they did cause the BBC to delay the broadcast by several weeks.
https://www.bbc.com/news/uk-67884743
It wasn't just a "glitch" -- it was also a PR campaign (quite successful up to now) that supporessed the voices of those affected.
If you want a summary and insight into the staggering scale of the injustice, this article from Private Eye magazine is worth reading:
https://www.private-eye.co.uk/pictures/special_reports/justi...
This BBC radio programme, started in 2020, also gives a lot of good information including details of how suspected sub-postmasters were questioned by the Post Office.
As you said this didn't just involve software issues to get that big, but I still wonder how all the devs that worked on this over the years could just shrug it off all these years. You can't tell me not a single dev got wind of these issues. Did they just immediately jump to "impossible I'm the most awesome person on this planet no way there's a problem these people all just stole money". I'd have sleepless nights going over the whole system in my mind trying to figure out where things went wrong. Is this some sort of dunning -kruger going on? I don't really consider myself a coding wizard or anything, but I think I know about enough to know I'll keep producing bugs til the day I die. Rust won't stop logic bugs from happening.
They told every single victim "you're the only one who's claiming there are issues".
"The Post Office is owned by the government, through the Department for Business, Energy and Industrial Strategy (BEIS) and UK Government Investments (UKGI), however, the Post Office Ltd Board has responsibility for the operations of the Post Office" [0]
And "The Department of Business Energy and Industrial Strategy holds government responsibility for postal affairs, including the Post Office" ... "the Post Office Ltd remains accountable to the government" [0]
[0] https://researchbriefings.files.parliament.uk/documents/CBP-...
"There is no direct evidence of her taking any money [...] She adamantly denies stealing. There is no CCTV evidence. There are no fingerprints or marked bank notes or anything of that kind. There is no evidence of her accumulating cash anywhere else or spending large sums of money or paying off debts, no evidence about her bank accounts at all. Nothing incriminating was found when her home was searched." (The only evidence was a shortfall of cash compared to what the Post Office’s Horizon computer system said should have been in the branch.) "Do you accept the prosecution case that there is ample evidence before you to establish that Horizon is a tried and tested system in use at thousands of post offices for several years, fundamentally robust and reliable?"
My word against yours wouldn't be enough to meet the standard of "beyond a reasonable doubt", but the Post Office's word backed up by a computer system? It seems that was convincing enough for the jury. They gave a guilty verdict in the above case.
Surely the system must be able to spit out what services and products were sold and how much they would be worth.
Many commentators are saying that this presumption should be changed:
https://www.theguardian.com/uk-news/2024/jan/12/update-law-o...
https://www.forbes.com/sites/emmawoollacott/2024/01/15/law-o...
I guess part of then problem is that the justice system takes every case in isolation, but the legal system really needs some mechanism where there’s a “hang on, something is wrong here” after the first few…
First change in this case specifically is probably stopping the archaic convention of the post office making their own prosecutions in the UK…
Looks like Wikipedia has termed this Algocracy (government by algorithm).
It was really a result of the Liberal & National Party (LNP) who were in government at the time wanting to:
- look tough on welfare cheats for image/election purposes
- claim a potential massive source of income (that wasn't really there) for the next budget
Most of the wrong calculations were due to a result of income averaging fortnight or multiple fortnights worth of income and extrapolating them to an annual figure, which is the absolute wrong thing to do when a lot of people on welfare worked as casuals and/or had short term jobs and the fortnight(/s) were non-representative of their yearly income.
In some cases the first time the alleged debtee found out about the debt was via a demand from a debt collection agency.
When individual action was taken against them, the government was reducing debts to $0.00 and then claiming there was no longer a reason for a case against them and had it dropped. The only reason the debt recovery program was stopped was because lawyers acting upon someones behalf also claimed interest on a paid false debt, which meant they couldn't succeed using the same tactic at which point they rolled over and admitted the unlawfulness of the scheme, which had been going for many years at this point.
I recommend this youtube 3 part series to get a good understanding of what was involved:
[0] https://www.youtube.com/watch?v=OfsL9GAbl3M
It starts out mentioning that it resulted in the largest class action lawsuit in Australian history...
Ironically the same reason for the Horizon system which ballooned into the PO Scandal.
What is needed is the requirement that software decisions must disclose their data and decisions path/ “algorithm” in court.
Another thing we need laws for is banning a person from using a system. It’s insane that you can be banned for life without recourse or explanation. It’s basically you being thrown in jail for life without reason.
In Denmark the police thoroughly investigates cases also when the accused plead guilty.
In a country that has a rule of law, punishment is not negotiation. I do realize that it is up for interpretation for the Anglo system due to its over-commercialized legal system.
In short, if the computer said so, it is a fact, in court.
Lots more information here
https://evidencecritical.systems/2022/06/30/briefing-presump...
This is kind of true, but there's a bit of nuance -- presumptions can be rebutted by the other side adducing evidence to the contrary. The evidence does not have to be strong enough to "prove" anything, it's generally sufficient that it suggests the presumption might be invalid. So in colloquial terms you don't have to prove the computer system was flawed, just a real possibility.
That said, the article you linked to does go on to explain why even this scheme didn't work out in the post office cases.
2x fairly prominent Fujitsu people have been interviewed under caution on suspicion of perjury, and some of the recent revelations from recent testimony in both the parliamentary committee hearing and the public inquiry seem likely to open the door to further such investigations.
In terms of internal conduct, the Post Office is now (as of first week of January) apparently under investigation for potential fraud offences. There is some speculation this includes individuals, as well as the entity itself (given many fraudulent actions would be carried out by individual people).
I imagine such prosecutions (properly carried out, independent of government actions) may help focus people's memories and also begin to start the process of punishing those individuals concerned, and implicate others in the process.
If a post office owed a billion pounds then that would be impossible to blame on the postmaster.
I don't think the "I was just trying to show how ridiculous the bug was" defense will go too well in court.
Given many of the sub post masters struggled to fund their legal cases (and legal aid is not likely going to stretch to funding the intensity of expert witness work required to discredit a whole system), this also seems to be one of the non technical failings not often talked about.
There's certainly questions about why courts didn't spot issues, but ultimately the common law adversarial system assumes both parties can get equal quality of representation. Unless you have very strong technical knowledge (this site is a non representative sample!) you'd struggle to really shed any light on the bug and get a court to agree with your line of defence.
Ultimately that's what did happen in the group litigation by Bates, about which the TV drama was ultimately created.
Lots of things had to go wrong for this miscarriage of justice to happen, but at the core of it all is the fact that the Post Office lied for years - to the defendants, to the courts, to the media, to government. Things should have been handled differently, the case has revealed a number of systematic failings that need to be urgently addressed, but there's only so much the legal system can do to defend against very sophisticated people who engage in a determined campaign of deceit. We can reduce the risk of something like this happening again, but sadly we can't eliminate that risk entirely.
https://davidallengreen.com/2024/01/how-the-legal-system-mad...
The problem that this unearthed was that evidence of crimes committed through information systems can be obscenely complex and therefore obscenely expensive to defend against.
"Computer says guilty" shouldn't be enough, but a defence would take months of debugging. Not something somebody on a £20k salary could ever afford.
But that's what happend. Hubris that Fujitsu's system was infallible. Targets and bonuses that stopped management asking uncomfortable questions. Layers of incompetence meaning people weren't asking the right questions, missing the correct burden of proof in the legal process.
And all this over an accounting system that can be forensically picked apart. Just imagine how bad it'll be when it's a black-box AI.
Its pretty much the biggest tech scandal in the news right now.
https://news.ycombinator.com/item?id=38937705
Actually the article has been posted since 2 years ago, but didn't get much discussion until recently.
https://hn.algolia.com/?q=https%3A%2F%2Fen.wikipedia.org%2Fw...
The hardest part is probably handling and recovering from all possible failure scenarios. You need to make sure that the system could crash while in the middle of processing any line of logic in your system and it should be able to recover elegantly; without skipping anything and without re-processing what has already been processed (which can cause duplication of records).
The challenge with distributed/partitioned systems specifically is that atomicity is much harder to achieve and strategies for achieving a similar result are complex and error-prone (e.g. two phase commits, using idempotency to avoid double-insertion)... For complex database transactions involving several tables with a custom two-phase commit mechanism, you have to be careful to process records of different types in a specific order. Also, you need to set up your database indexes carefully for fast lookup and sorting...
You get the software to print off a list of every sale made that day with timestamps. You check the stores CCTV so see when every customer came into and out of the store. You cross reference. Most are small stores that probably only see 300 customers all day anyway.
If you are a store manager wrongly accused of fraud, of course you can be bothered to do this.
It would only take one store to do this to find out that the software was wrong.
Typically, when it gets to the point of being an accusation, employees will lose all access to all systems. It's also widely reported that evidence was routinely suppressed and emails conveniently lost.
And the software was found wrong in November 2019, when a judge ruled that the software contained bugs, errors and defects, which caused a settlement in that case and a cessation of further cases.
The oddest part is that with all these Horizon problems being pointed out for decades and then a judge ruling that it was faulty, there still hadn't been an inquiry until very recently.
It's wholly owned by the Government and there is a minister responsible for postal affairs, which includes the Post Office. However it has its own board for day-to-day operations.
The government has a _lot_ of blood on their hands
Source: https://researchbriefings.files.parliament.uk/documents/CBP-...
It was a software glitch, but the coverup was massive, even years ago. It was intentional to hide their incompetence.
Absolutely no they wouldn't.
"Mistake? Ha! We don't make mistakes."
Proceeds to drop ceiling plug through ceiling.
"That's bloody typical. They've gone back to metric without telling us."
It was corrupt managers, police and goverment! Literal conspiracy to destroy lives of innocent people!
The comment I replied to literally said exactly that.
The organisation in the public sector must choose the right vendor and then any sh*t sold by this vendor is correct, otherwise it would prove the wrong choice.
P2: oh Im responsible for the deaths of many people and court cases spanning 2 decades. There's even a tv show based on the story
It was people making decisions about suppressing evidence, lying, not correcting the mistake.
It’s like saying faulty breaks killed thousand people and omitting that the company could have recalled the car after the first car but chose not to.
Is suspect this kind of reporting is intentional to minimize pushback. “How Fujitsu and the Post Office drove people into suicide” gets you angry lawyers. Blaming the glitch not.
It wasn't a "software glitch" that caused anything. Yes the software sucked, but lots of software sucks. It was human decisions that caused this, by using software they knew (or could and should have known) was not fit for purpose. And by ignoring people who told them it was not fit for purpose after it was deployed. And the government not stepping in. And National Federation of Sub-postmasters. &c &c &c. All of this is well documented.
That the initial mistakes were made was kind of ridiculous for lots of reasons, but okay, mistakes happen. But that it took 20 fucking years to correct while the broken software continued to be used is just indescribable, and absolutely not the fault of any software but the result of human choices. That's the real problem; in an alternative universe people realized Horizon was a piece of crap shortly after being deployed, the mistakes were corrected, and they fixed it or stopped using it.
Humans caused the misery. Not software. The software bit is almost a minor detail IMO. By framing it as "a software glitch" people will focus on software to "fix" things, but that's not where things need to be fixed.
100% agree. Managers who were in charge of these choices should be convicted.
They should have the book thrown at them.
There is plenty of strong evidence that this is the case. Stronger than the evidence used against the post workers accused. Take note that no one is threatening the show runners with libel actions & such…
> They should have the book thrown at them.
And thrown hard. This is not a set of honest mistakes, but a protracted campaign to pervert the course of justice.
If nothing else this puts into focus plenty of bureaucratic problems, but we wouldn't be here without the underlying buggy software that opened the door for bureaucracy to fail so badly.
Yes, because as the OP clearly stated, the problem wasn't the bugs.
The problem was the bugs were noticed, flagged, discussed, _and systemically ignored_.
Mistakes get made. And even if the original mistake was hugely egregious, it's usually not a huge problem if you actually deal with it well (barring things like medical errors, but even there, trying to cover up a mistake can cause more harm than the initial mistake, at least in some cases).
For these kind of institutional problems most of the time it's really the response to the mistake that turns a mistake in to a huge catastrofuck.
describing this as a software glitch is kind of heinous in its own right.
Your game crashing is a software glitch, what happened in the UK is far beyond that. It'd be like describing the issues Boeing is experiencing as a "hardware glitch". Decisions were made that should not have been made, full-stop.
It wasn't a "bridge collapse" that caused anything. Yes the bridge sucked, but lots of bridges suck. It was human decisions that caused this, by using a bridge they knew (or could and should have known) was not fit for purpose. And by ignoring people who told them it was not fit for purpose after it was deployed. And the government not stepping in. And Association of Professional Engineers. &c &c &c. All of this is well documented.
That the initial mistakes were made was kind of ridiculous for lots of reasons, but okay, mistakes happen. But that it took 20 fucking years to correct while the faulty bridge continued to be used is just indescribable, and absolutely not the fault of any bridge but the result of human choices. That's the real problem; in an alternative universe people realized the bridge was a piece of crap shortly after being deployed, the mistakes were corrected, and they fixed it or stopped using it.
Humans caused the misery. Not the bridge. The bridge bit is almost a minor detail IMO. By framing it as "a bridge collapse" people will focus on bridge design to "fix" things, but that's not where things need to be fixed.
When will the software sector be regulated like other sectors?
But you just said the SW sector didn't cause this, it was those who ordered it and were responsible for managing the public service using it.
20 years and a high-profile TV serialisation heightening public perception of the problem a little too close to an election for comfort. The government is running around saying they care and how much of a tragedy they think it is not because they genuinely care but because the voting public currently care and there is a general election within the next 12 months.
One of the massively empathy lacking comments from a couple of weeks ago (prior to the proposed law change specifically to speed up exonerations for this scandal, announced last week) was that they couldn't just quash all the convictions without retrial because that might let some genuine cases of fraud through – as if letting one or two criminals off more lightly (they have had 20 year of whatever punishment was originally handed down and its side effects) is such a problem compared to continuing to punish hundreds of innocents. Let them all off, then retry those few if you find new compelling evidence. But before doing that: get the people who tried to cover it all up, who bullied many into silence despite knowing (at least some of) the truth, and those who just stood by and let it all keep happening, etc., both in the post office and the government, into court first – what they have done is far more reprehensible than what those post workers were (wrongly) accused of. Handing back one CBE just doesn't cut it.
Meanwhile, if I want a multi-billion dollar critical infrastructure software project, I can use people who learnt to code at summer camp. Isn't it time we had proper qualifications for software 'engineers'?