A second fun one.
In a previous life, I was paid to build storage systems for some HPC-ish workloads.
So I was testing a bunch of cheap-ish desktop drives that had really nice (for the time) sequential throughput, supposedly, and threw together a Supermicro system with a Xeon, ECC RAM and some SAS HBAs and several external enclosures full of these disks, made a couple of raidz3s, and started trying to stress it.
I quickly found that rarely but somewhat reliably, I'd get correctable checksum errors, and because it was raidz3, it always had a lot of spare recovery bits, but the numbers kept going up...and went up even after a scrub finished and "corrected" all of them.
Well, SAS has checksums over the wire, so I'd be seeing disk errors if the wires were eating my bits, and it was across all the disk controllers, so either they were all bad or it wasn't the controllers, ECC RAM and no correctable or uncorrectable events fired...
This predated ZoL, so this was originally on early illumos - I then tried FreeBSD, and it did the same thing.
Huh.
So, the drives in question were Samsung HD204UIs, which, it turns out, have a really spicy firmware bug, where if you send them a SMART IDENTIFY request with data in the write cache (e.g. it already told the OS it was stably written out), it just...dropped it.
So the background smartd I had running for collecting data on the disks was causing them to eat some of the writes whenever that lined up.
The funniest part was, though, that Samsung released a firmware update, but it doesn't change the reported firmware revision, so you can only know by testing if the disk is going to eat your data like that.
I believe smartctl still prints a big warning to this day about all this if you ask it about those drive models.