Please Use ZFS with ECC Memory (2014)
louwrentius.com
louwrentius.com
Noticeably more stable, as in, "never ever crashes" as opposed to "almost never crashes" without ECC. Thanks Linux! :D
Memory errors in data centers tend to be concentrated in a small number of bad sticks of RAM rather than evenly distributed across all memory. If you have a machine crashing regularly due to memory errors, it’s likely a bad stick of RAM, not random errors due to lack of ECC.
Uncorrectable errors will cause a machine check, unless the BIOS disables it. Which some do.
You are absolutely correct that reporting of errors may not be "honest", I prefer to use the work "correct" rather than honest because honesty implies and intention to lie and the computer logic is generally designed to be correct.
The market reality is that reporting correctable errors can generate unnecessary service tickets as people who are not technically sophisticated. Those users may generate a support ticket wondering why their system is telling them it saw and error but then it corrected it. Dealing with support tickets costs money, money that is taken away from margin, and so it is sometimes "optimized out" which is that some manager makes their numbers better by ordering the software team to "hide" correctable errors which were corrected.
It is absolutely fair to describe that management choice as "dishonest."
It can be difficult sometimes to find out where the dishonesty was applied however. It has been my experience that between the BIOS and the kernel, if the chipset supports ECC and the BIOS recognizes that you have ECC memory installed, there is a means for extracting all ECC events reliably. It isn't always well documented and sometimes requires several levels of escalation in support to get the information you need, but when ECC is important to you it can be worth it. It can also inform your choices for vendors to use in the future. :-)
My ECC workstation had lots of memory issues where it blue-screened on the windows side. For some reason that wouldn't happen on the Linux side.
Many errors won't cause a crash, but said errors can accumulate in processes, filesystems, and files and you won't know why. 10 years from now you may find a photo, music file, or binary that's somehow corrupt and have no idea how it happened.
Interestingly, this never happened when I was running Linux and abusing the RAM with photogrammetry workloads. It would only happen with windows when I wasn't using anything beyond routine levels of RAM.
> There's nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem. If you use UFS, EXT, NTFS, btrfs, etc without ECC RAM, you are just as much at risk as if you used ZFS without ECC RAM. I would simply say: if you love your data, use ECC RAM. Additionally, use a filesystem that checksums your data, such as ZFS.
And goes on to say
> I have nothing to substantiate this, but my thinking is that since ZFS is a way more advanced and complex file system, it may be more susceptible to the adverse effects of bad memory, compared to legacy file systems.
The truth is that everyone should want to use ECC wherever data integrity is a priority, regardless of whether or not they’re using ZFS. The conjecture about ZFS being uniquely vulnerable has been debunked over and over again. Maybe there was something different back in 2014 when this was written, but even if that was the case it’s not a good idea to use such old blog posts full of admitted conjecture to anchor your technology understandings.
If you're Google (you're not Google) then the scale of errors becomes a choice of optimization.
It's some sort of variation of the Toupee Fallacy, you aren't aware of the frequency of bit errors because you have absolutely no way of knowing when they happen, and what "random" errors were caused by them.
I don’t care about triggering an occasional bug or a bit flip in a video stream, these things are inconsequential. If a bit gets flipped in a photo it’ll be a very slightly different photo. A piece of text will be slightly different… if either had the luck to be in memory ready to write when hit, in dozens of gigabytes of memory.
The most likely consequence of a bit flip is nothing at all happening because it hit unallocated memory. The next most likely consequence is nothing as it hit a piece of program memory that has no effect on operation and gets overwritten in moments.
ECC zealots are really excited about preventing something that just doesn’t happen all that often: actual data corruption with permanent significant effects.
This is really weird. How come even routers have ECC memory while our actual computers don't?
Why can't all memory just be ECC memory? I don't even see the point of non-ECC memory. Everything has error correction built-in: hard drives, SSDs, even optical media. Yet RAM doesn't?
(The 1st generation of Pentium, the 60/66 MHz models, had been introduced in 1993, while the 2nd generation, the 90/100 MHz models, introduced in 1994, used a different socket and different motherboards.)
Unlike with the previous Intel CPUs, Intel had the idea to introduce a market segmentation between Pentium and Pentium Pro (for the successors of Pentium Pro, Intel created the Xeon brand, a couple of years later).
Therefore, Intel has retained the ability to detect and or correct memory errors only in the Pentium Pro chipsets (at that time the memory controllers were in the so called northbridge, a part of the chipset, and not in the CPU), while removing this feature from the Triton chipset intended for Pentium.
While Intel has initiated this market segmentation policy, between "amateurs" and "professionals" (unlike IBM, which since the first IBM PC had always included memory error detection via parity), both the memory module vendors and the makers of other CPUs have been also happy to follow this policy, because it allows them to extract much more money from the knowledgeable users than the cost of adding ECC, while also having larger profit margins when selling to naive users, because the elimination of ECC/parity has never caused any price reduction.
When this feature was removed, all the new motherboards and modules without parity/ECC had the same prices as the previous models with error detection (for many years, the memory errors were detected via parity, but then the memory controllers were improved, so that when using the same number of extra bits, i.e. the same memory modules and motherboards as for parity, single errors could be corrected, not only detected).
The frequency of DRAM errors is proportional with the amount of memory installed in a computer. For smaller quantities of memory, e.g. 16 GB or 32 GB, the frequency might be of one error every few months. With ECC, the mean time between errors is likely to become longer than the lifetime of the computer.
Intel, who has initiated this absurdity of convincing the naive customers that it is OK to buy computers that may make mistakes from time to time, has succeeded to get away with it only due to the huge amount of software bugs, which have made the computer users believe that whenever a computer crashes, or some corrupted data is discovered, it is more likely that the cause has been a software bug and not a hardware defect.
The integrated circuits made with modern CMOS processes, with very small devices, have a non-negligible aging rate.
Because of that, memory modules that have been used for many years start at some point to have much more frequent errors than when new.
When you have ECC, such memory modules are immediately detected and they can be replaced before causing some irremediable data corruption. This helped me a lot a few times.
I think the misunderstanding goes something like this. I'm pretty sure all three of the below statements are true:
• People who prioritize data integrity are more likely to be using ZFS.
• People who prioritize data integrity should use ECC memory.
• Ergo, if you are using ZFS, you should probably be using ECC memory.
However, it doesn't naturally follow that because you are using ZFS, you should be using ECC memory, nor that people without ECC memory should not use ZFS. This is unintuitive.
I use a usb-attached zfs raid-1 for my important storage. I have to run "zfs-fuse" before I mount it.
What I know about ecc comes from this:
https://arstechnica.com/civis/viewtopic.php?f=2&t=1235679&p=...
I have never scrubbed my zfs vault, but I have scrubbed my btrfs drives.
I should get something with ecc. Wow, do I have a lot of stuff.
I have no idea why one would use a 2006 Google summer of code project instead of the official openzfs project which provides an actual kernel module that would be several times faster and likely safer.
The problem was the same even in 2014: ZFS lets you know how shitty your hardware is.
And, just like 2014, people would rather shoot the messenger because they don't have any power to force companies to make better hardware.
The only things that really happened was that a) ZFS exposes how failure prone hardware is due to pervasive checksumming b) ZFS came from culture (and project goals) that prioritized data safety - as such, it was possibly first contact for many with strong recommendations about ECC memory
The amount of people rehashing this "you must use ECC with ZFS" was insane.
One guy seemed to have made it his life's mission to publish this same story over and over again.
ECC is corruption of the parts in memory, like caches of data, data when it's in the process of being compressed or encrypted, some management data always cached in memory etc.
In the end not having ECC can affect your system in such abstruse ways that IMHO it probably doesn't make any relevant difference if you use zfs or e.g. ext4.
Tbh it's kind absurd that ECC is not the standard for any non-"cheap" PC/Laptop.
EDIT: Checksums can still potentially help in the unlikely case that the RAM error is a bitflip in some loaded offset/ptr referring "into" the HDD/SSD storage (e.g. position of some file on disk).
No, it won't.
https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
> Let’s assume that we have RAM that not only isn’t working 100% properly, but is actively goddamn evil and trying its naive but enthusiastic best to specifically kill your data during a scrub. First, you read a block. This block is good. It is perfectly good data written to a perfectly good disk with a perfectly matching checksum. But that block is read into evil RAM, and the evil RAM flips some bits. Perhaps those bits are in the data itself, or perhaps those bits are in the checksum. Either way, your perfectly good block now does not appear to match its checksum, and since we’re scrubbing, ZFS will attempt to actually repair the “bad” block on disk. Uh-oh! What now?
> Next, you read [a copy of the same block from another disk]. Now, if your evil RAM leaves this block alone, ZFS will see that the second copy matches its checksum, and so it will overwrite the first block with the same data it had originally – no data was lost here, just a few wasted disk cycles. OK. But what if your evil RAM flips a bit in the second copy? Since it doesn’t match the checksum either, ZFS doesn’t overwrite anything. It logs an unrecoverable data error for that block, and leaves both copies untouched on disk.
If there is a filesystem that is dumb enough to cause corruption during the checksumming process, please let me know which one, so I can be sure to never ever ever go anywhere near it. :)
Error checking will only ever help you, not hurt you. It doesn’t matter how bad you memory or disk or raid controller is. Error checking won't necessarily save you from those things, but it can in some cases, and it’ll never make things worse.
"There's nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem. If you use UFS, EXT, NTFS, btrfs, etc without ECC RAM, you are just as much at risk as if you used ZFS without ECC RAM. I would simply say: if you love your data, use ECC RAM. Additionally, use a filesystem that checksums your data, such as ZFS."
Wow. If you weren't talking about ZFS in a comment under an article concerning ZFS, an article concerning the same issue you are parroting, an issue that supposedly effects ZFS, what filesystem were you talking about?
Let me guess, a hypothetical filesystem.
Sure, whatever.
> bad memory corrupting data being actively changed remains a problem for any filesystem. If the filesystem is actively changing data it can not rely on anything it is writing to disk being correct if the in memory buffers and other data structures are themselves corrupted.
And, yet, that's not the argument you just made. This is what makes me thinks this is a bad faith pivot to something, anything respectable. It's not like I'm going to forget that you said right above (it's still there!) that a scrub could destroy all your data. This claim has been repeatedly debunked. You may have been taken in. Fine. That happens. Just, now that you know, don't keep spreading this FUD.
It's as suspect as I have nothing to substantiate it but my kid caught autism after that shot.
It's so badly reasoned it calls into question the authors judgment on any technical topic because it says his basic tools for reasoning are lacking.
99.9% of data loss occurs because of well understood but difficult to fully mitigate problems.
There is no reason to believe zfs is more susceptible to bad memory and it's not exactly new having begun development 21 years ago.
And zfs uses more and relies on it more than others, above that baseline.
Sure, the corruption can occur, but if the memory is not faulty, ZFS will detect data corruption in-situ during a scrub. It's just that ECC errors can introduce their own corruption that is separate, during write operations only.
I believe if you have no choice but to use non-ECC (e.g. existing hardware that is limited by Intel's stupid design choice of "ECC is only for servers"), ZFS is still much better than using ext4. It still protects against a different class of errors which is HDD degradation. When used in RAID-Z it can even recover for them.
For perfect protection ECC is necessary. But this is not always possible financially. I think it's a bit of a wild statement to say that if you can't afford a server with ECC you should forget about ZFS entirely.
I would rather run XFS or EXT4 with ECC memory than ZFS without ECC because silent bitrot is extremely rare because drives and protocols are full of error correction stuff.
Memory is the real weak spot from a bitrot perspective.
ECC first. ZFS second.
ZFS does not have such [recovery] tools, if the pool is corrupt, all data must
be considered lost, there is no option for recovery.
In other words, this marvelous piece of over engineered technology is fragile. When it works, it's great. When it fails, it's a complete disaster.Everything fails eventually. How do you prefer your failure served? In small manageable increments or one spectacular, complete, overwhelming, unrecoverable helping.
Multiple rotating backups would have solved your problem before ZFS.
Also, on a non-checksummed filesystem, how will you know whether a file needs to be restored from the backup?
Has same edit epoch? If yes skip. If no compute checksum of both and compare, copy over in case they differ.
I am fully aware of the flaws of this.
zpool import -FX mypool, where:
1) -F Attempt rewind if necessary.
2) -X Turn on extreme rewind.
3) -T Specify a starting txg to use for import.
There are a few other points you might consider:First, what kind of failure might cause an unrecoverable corrupted pool, and how useful are other filesystems data recovery tools actually, if corrupted in a similar fashion? Is there an apples to apples comparison that you can share? I believe the author and you are sharing speculation.
Second, how do you know you have corruption with another filesystem until you read back the data? And do you even know you have corruption then?
Usually bug => corruption => module crashes => no zfs own tools work.
Second, you managed to ignore my two other original points -- 1) Is the state of other filesystems any better re: recovery in similar circumstances (show your work, when XFS is afflicted with the exact same type of corruption, how is it better)?, and 2) is it possible are you better positioned to deal with corruption with ZFS (because it's usually not silent)?
I'm not saying ZFS is better for you. You are obviously not predisposed. What I am saying is -- your argument has to be better supported for it to make any sense.
Maybe such an absolute statement the author made is somewhat incorrect, but it's not entirely incorrect. Empirically so, according to the issue tracker. If people didn't request these things every year, I'd agree that there are no recovery issues.
I do think that other filesystems tend to have better 1st and 2nd party tools for repair and recovery. It is also true that some file systems are less complex and that might help. But isn't that an argument for providing even stronger 1st party tooling and documentation? We want people to not lose any data, no?
So to clarify the discussion a bit, what is it that you're claiming? Are you claiming ZFS has no recovery issues or that it's okay because some others are equally terrible to recover?
And I think there is a kooky caucus for half ass solutions. The Linux community has chosen to beclown itself by pretending ZFS isn't absolutely amazing (and free!), and the result is the strangest collection of system level kludges anywhere.
You forgot Stratis and unraid!
Well, there is a huge number of reasons for that.
Just recently ZFS got RAID-Z expansion capabilities. Before that, I would have had to make plans for a storage server and buy all the storage devices upfront. Now I can expand the storage pool as needed, one device at a time.
The ability to do this is the only reason I even bothered with btrfs. Now that ZFS has flexible storage pool expansion, it is essentially perfect as far as I'm concerned.
The goal is to never reach this point, via a combination of redundancy and backups. ZFS helps enormously with the first part.
It's not as black and white as you or the blog post makes it out to be. I've had to recover damaged ZFS pools while traveling (there are neither affordable nor travel compatible no laptops with ECC RAM support). ZFS scrub told me which 4 files have been corrupted and overwriting them with good copies from an other machine solved the problem. Without ZFS (or a similar file system) I wouldn't have noticed the data corruption as quickly and wouldn't have known what to restore. Also ZFS is no one trick pony and has more to offer than "just" end to end checksumming e.g. pooled storage for multiple file systems and sparse block devices, fast consistent snapshots (good enough to backup a running RDBMs), incremental backups and replication (no more dreaded full backups), ease of administration, easy to grow capacity (as long as you do it in large enough increments), transparent compression and if you really need it block level deduplication (you probably don't want online block level dedup).
I personally am glad that ZFS doesn’t have “easy to partially recover a corrupted FS” on their list of priorities. Designing for that might take resources away from making the system as a whole more robust, and at best it’s a half-measure that will never be a substitute for an offsite backup solution.
Backups have a delay. A filesystem that fails catastrophically is significantly worse than the one that doesn't.
I really do not see how anyone could possibly think that is worse than your data being silently corrupted on disk, which then propagates to your backups, corrupting them too.
People seem to think ZFS is worse because it tells you when your data is corrupted, rather than not telling you? It doesn't make any sense.
In the best case, when good data was written to disk, checksumming preserves the good data. In the worst case, when bad data was written to disk, checksumming does no harm.
Checksumming is good.
It is true because ZFS on Linux crashes very quickly when metadata is corrupt. It can't roll back transactions because it'll crash almost instantly even looking at your pool, not to mention when scrubbing it.
The only time I’ve had ZFS fail, was when both drives in a mirror pool died within 8 minutes of another. No filesystem would save you there.
Since ZFS has a replication protocol (ZFS send) recovery was basically a few ZFS revc calls into a new pool.
Running ZFS on a home server as a NAS but without ECC memory, you are addressing one risk, but you are forgetting another (IMHO bigger) risk: faulty memory causing bitrot.
It's as if you lock your back-door and leave the front-door unlocked.
Also, bitrot is not like burglars. Burglars try one door and if they don't succeed they will look for further vulnerabilities. Memory and disk corruption is random, it's not trying to attack you.
In many cases ECC is just not possible without buying expensive new hardware (with possibly higher power consumption) due to intel being so precious with ECC as a 'server feature only'. Data protection is not absolute, protecting against one class of corruption is better than none. E.g. one of my NASs is a low-power NUC specifically chosen for its power consumption.
Of course both is even better but not everyone has the financial resources for perfection.
In the end people worry too much, all consumer NAS gear doesn’t have any bitrot protection… (to put things in perspective).
Any recommendations for what to look on ebay?
My NAS is a 2U Dell R510 (Dual E5649, 32gb ECC memory), I think it cost me about $250 on Craigslist, and I've had zero issues. It will likely be difficult to find a machine that can take 10 drives, mine can take 8.
If a machine has 3.5" SAS bays, it likely can take SATA drives as well, but be sure to do your own research. Also, in a lot of these older machines the PCIe card that connects the SAS backplane is hardware RAID only, which means you will have to find a new card, and possibly flash it with a new firmware to enable JBOD mode, so that your disks are passed directly to the host OS. I had to do this for my NAS but it wasn't difficult.
One other tip I highly recommend, you can buy PCIe cards that fit one or two 2.5" SATA SSDs in. I run my TrueNAS off of a mirror of two small SATA SSDs this way, which frees up space for more hard drives in the front.
But you can activate Arc check-summing at least.
That's a better article:
https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
My favourite aspect about ZFS is the 13yo Python script for reverting transactions by destroying uberblocks being the to-go even now.
Edit: sorry, this is misleading. See below or ignore. Ddr5 "ecc" that I am referring to is not the same as what one would think when saying ecc normally
So this is one more thing that will confuse the consumer.
What does consumer DDR5 do when there's a correctable error, what does it do when there's an uncorrectable error?