https://www.youtube.com/watch?v=vxFNBZIAClc
and they could not make any errors. This was pretty brutal. When I saw this video, I decided I never want to use any other filesystem than ZFS ever.
https://www.youtube.com/watch?v=vxFNBZIAClc
and they could not make any errors. This was pretty brutal. When I saw this video, I decided I never want to use any other filesystem than ZFS ever.
Overclocking some components might work too.
I mean, if the ship is on fire, you want the missile control systems to shut down after all missiles have been fired.
The best part about ECC RAM, IMO, isn't the correction, but the checking. I want to know when my RAM goes bad, and I'd prefer to know before ZFS detects a problem.
This is by the inventor, Matthew Ahrens: “There’s nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem.”
https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
Now give me my karma back.
Furthermore, there is error correction at the chip level. I think a fitting analogy is like this: A single-engine airplane might seem much more dangerous than a dual engine airplane, but most single-engine airplanes have dual fuel pumps, dual ignitions sources, and additional auxiliary devices sharing the same engine body. Will it ever be as safe as a dual engine plane, no, but it's not as dangerous as having no failsafes.
Modern RAM is believed, with much justification, to be reliable, and error-detecting RAM has largely fallen out of use for non-critical applications. By the mid-1990s, most DRAM had dropped parity checking as manufacturers felt confident that it was no longer necessary.
"The SDRAM and DDR modules that replaced the earlier types are usually available either without error-checking or with ECC (full correction, not just parity)."
ECC is necessary for anything real or important based on the financial cost of losing data. For everything else, it can be optional.
Plus, don't forget the security issues of bitsquatting and other attacks that are the real results of silent bitflips.
DEF CON 19 - Artem Dinaburg - Bit-squatting: DNS Hijacking Without Exploitation https://youtu.be/9WcHsT97suU
Non-ECC - Checksum: parity bit BlockSize: 1 byte
ECC - Checksum: 7 bits BlockSize: 8 bytes
ZFS - Checksum: 256 bits BlockSize: ~128kb
The issues are more subtle beyond hardware though - your OS kernel has to understand the detection from the BIOS / EFI and if the motherboard manufacturer decided not to opt for wiring in the checksum fail signal you will run blind to data corruption that’s potentially really subtle on typical consumer hardware. With some BIOSes RAM failures result in a hard panic and will hard reboot the machine (the idea being a crash is better than continuing with an error).
For me, ECC (both UDIMM and RDIMM) today are so inexpensive it’s a no-brainer for builds for business.
There is no source checksum.
> transmitted file's checksum will not match the source's checksum
Most applications don't checksum their files when reading them.
> a retry will likely solve the issue
You can't retry; the only copy of the edited file is corrupt, and it's not possible to (automatically, in the general case) distinguish edits you made since loading the file out of ZFS from memory corruption that happened since loading the file out of ZFS.
2. If an error occurs on ZFS the consequences are worse, because there are very limited recovery-tools available for ZFS. The official checkdisk is "just restore from tape, it's quicker than fsck anyway". Very enterprise oriented.
Now you might be okay with the above. But the risk is greater and the potential consequences are worse. Not because of ZFS being special, but because of the way it makes use of your hardware.
Also keep in mind that a lot of people pick ZFS to reduce potential corruption so the addition of ECC is par for the course.
well not exactly true. we know that the raid5 write hole won’t tolerate random power off / memory removal etc.
as someone else said, x ray test would have been better.