Computer chips are mercurial: Rare miscalculations frequent at cloud hyperscale
theregister.com
theregister.com
That's the typical bit error rate quoted by manufacturers of network devices and hard drives. Physically, something like a PCIe bus is also a network link, and has similar bit error rates.
Of course, some of these are detectable, and some are correctable. But some are uncorrectable, and the worst ones are undetectable.
In my experience, digital signing or encryption for protocols helps noticeably reduce random crashes and weird unreproducible errors. Similarly, disk formats like ZFS, BTRFS, and ReFS that keep cryptographic hashes of the data also regularly report corruption when used at scale.
A rule that I've come up with is:
Security is synonymous with robustness.
Things that improve one almost always improve the other, and neither is really possible without the other. Cryptography is not just there to keep the hackers out, it also ensures that bit flips are caught before they crash your mission-critical application.
There's certainly a large overlap. At the same time you can use insecure hashes for robustness, so the implication doesn't go both ways. That's why I would say security usually implies robustness.
> On Sanjay’s monitor, a thick column of 1s and 0s appeared, each row representing an indexed word. Sanjay pointed: a digit that should have been a 0 was a 1. When Jeff and Sanjay put all the missorted words together, they saw a pattern—the same sort of glitch in every word. Their machines’ memory chips had somehow been corrupted.
> Sanjay looked at Jeff. For months, Google had been experiencing an increasing number of hardware failures. The problem was that, as Google grew, its computing infrastructure also expanded. Computer hardware rarely failed, until you had enough of it—then it failed all the time. Wires wore down, hard drives fell apart, motherboards overheated. Many machines never worked in the first place; some would unaccountably grow slower. Strange environmental factors came into play. When a supernova explodes, the blast wave creates high-energy particles that scatter in every direction; scientists believe there is a minute chance that one of the errant particles, known as a cosmic ray, can hit a computer chip on Earth, flipping a 0 to a 1. The world’s most robust computer systems, at nasa, financial firms, and the like, used special hardware that could tolerate single bit-flips. But Google, which was still operating like a startup, bought cheaper computers that lacked that feature. The company had reached an inflection point. Its computing cluster had grown so big that even unlikely hardware failures were inevitable.
https://www.newyorker.com/magazine/2018/12/10/the-friendship...
Is it RAM bitflips, cosmic rays or other things causing a CPU to malfunction, like here, or is there a more mundane explanation for most freezes?
It can even be a problem with a single machine depending on the environment. Like I was messing around with a simple circuit on a bread board. One of the ICs had a schimit trigger and just the EMI from being on a bread board consistently would cause my desktop computer near by to lock up.
There is all sort of spurious noise from cosmic to man made that has the potential to screw up a chips operation
Heck even a dip in power for a few micro seconds could cause issues, and most UPS use a mechanical relay that takes milliseconds to kick in ect... All though any decent PSU should have enough filtering capacitance to filter out such a short event
The failure rate and failure modes vary from model to model and even from batch to batch. The biggest issue is not that failures exist, but that they're hard to trigger in testing and when they trigger they do not cause a crash. The CPUs that any actually used computation is running on have already been through a lot of testing (manufacturer, cloud provider), the failures have to slip through all of them.
And the failures can be really nasty: imagine some distributed DB server damaging large % of database metadata records because 1 instruction malfunctioned, but didn't crash the process. By the time you know something's wrong you cannot recover 70% of the data because you have no idea which data blocks correspond to which logical rows (remember, the metadata is garbled), except through a manual process that can take weeks or months. Or something like a v-table "miss" where instead of something like Table.Info() the CPU calls Table.Drop() because that function happens to be exactly 64 bytes lower in the v-table and has similar enough signature for the call to succeed. Those are two real examples.
The only way forward is expanded testing, that's what the paper from Google is (also) about. I think this issue will always be with us, to a larger or smaller extent.
There's probably a ton of data corruption out there that happens to be in places that doesn't really cause big problems.
The problem is that nobody knows how to write tests that catch all errors.