But the root cause was bugs. So many bugs.
There isn't much the average HN poster can do about the political and justice problems, which are firmly in the realm of the British government. But there are people here who work on databases and app frameworks. What can be learned from the Horizon scandal? Unfortunately there doesn't seem to be much discussion of this. Compare vs the airline industry where failures are aggressively root caused.
I'll start:
1. Transaction anomalies can end lives. Should popular RDBMS engines really default to non-serializability by default (non repeatable reads, for example).
2. Offline is very hard. A lot of bugs happened due to trying to make Horizon v1 work with flaky or very slow connections, and losing transactional consistency as a result. The SOTA here has barely advanced since the 90s, instead the industry has just given up on this and now accepts that every so often there'll be massive outages that cause parts of the economy to just shut down for a few hours when the SPOF fails. Should there be more focus on how to handle flaky connectivity in mission-critical apps safely?
3. What's the right way to ensure rock-solid accountability around critical databases, given that bugs are inevitable and data corruption must sometimes be manually fixed? A lot of the Horizon problems seemed to involve Fujitsu manually logging in to post offices and "fixing" the results of bugs, in such a way that they didn't realize their fix created ledger imbalances that the SPMs would be blamed for. A part of why big enterprises got so excited about blockchains was this notion of an immutable ledger in which business records can't go magically changing around you without anyone knowing how. There are clearly ways to do this, but they're not the default.
4. IIRC at least some failures were traced back to broken touch screens generating false random touches, which could lead at night to random transactions being entered and confirmed when nobody was around. Are modern capacitative touch screens immune to this failure mode? If not, are consoles in embedded applications always reliably engaging screen locks?
I guess there are bazillions more you could come up with.