A broken memory module hid in plain sight
chollinger.com
chollinger.com
[1] https://www.intel.co.uk/content/www/uk/en/processors/xeon/xe...
It got used a lot to soak tests disk for servers that were going into production or being shipped out to far away places.
However it quite soon became apparent that a lot of the errors weren't caused by bad disks at all, but rather by bad RAM. It even discovered a set of bad RAM for 200 computers which passed all the manufacturers tests but failed when in use.
Moral of the story: if you care about data integrity check your RAM as well as your disks! I run memtest86 then stressdisk on any new systems I'm building.
Also, testing RAM is hard! Memtest86 does an excellent job, much better than the manufacturers tests.
Proper burn-in testing helps too; but ECC will help me later after things have aged. Maybe I should add a yearly re-validation to systems.
My current build is a borosilicate water-cooled dual Rome EPYC with ½ a TiB of ECC RAM. I already had it planned before the media went gaga over AMD. I just hope prices don't spike too much before I drop what's already going to be $13k to get this done.
PS: 72-bit (64,8) SECDED ECC RDRAM typically uses 4-bit-wide chips in multiples of 18 chips (# = ranks, 36 for 2R), whereas consumer 64-bit DDR4 uses 8-bit chips, in multiples of 8 chips. Also, does anyone remember IBM's Chipkill?
And get angry every time The Cloud serves us a painful reminded that it's just Someone Else's Hardware.
My favorite Faulty RAM story involves working for a hosting company that at some point homebrewed some Linux-based TCP/IP load balancers to offer customers instead of a proper F5 or whatever. This particular load balancer is restarting itself every time one of the Windows servers behind it fails a health check. Obviously not a desirable situation.
A tenacious young tech not yet battle-hardened to the fallibility of hardware had been working on the load balancer for hours. Eventually throws his hands up and calls for a Windows guy 'cause obviously Windows must be doing something weird to make the Linux crash. Linux never crashes, amiright?
Grizzled veteran that I am, I'm telling him to replace the RAM before he can finish explaining.
If you can't fathom what the problem could be, if it really isn't DNS[0], it's probably going to be RAM. If it's not RAM... update your resume and burn everything to the ground.