Battle testing data integrity verification with ZFS and Btrfs
unixsheikh.com
unixsheikh.com
https://www.youtube.com/watch?v=vxFNBZIAClc
and they could not make any errors. This was pretty brutal. When I saw this video, I decided I never want to use any other filesystem than ZFS ever.
Overclocking some components might work too.
The best part about ECC RAM, IMO, isn't the correction, but the checking. I want to know when my RAM goes bad, and I'd prefer to know before ZFS detects a problem.
This is by the inventor, Matthew Ahrens: “There’s nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem.”
https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
Now give me my karma back.
ECC is necessary for anything real or important based on the financial cost of losing data. For everything else, it can be optional.
Plus, don't forget the security issues of bitsquatting and other attacks that are the real results of silent bitflips.
DEF CON 19 - Artem Dinaburg - Bit-squatting: DNS Hijacking Without Exploitation https://youtu.be/9WcHsT97suU
Non-ECC - Checksum: parity bit BlockSize: 1 byte
ECC - Checksum: 7 bits BlockSize: 8 bytes
ZFS - Checksum: 256 bits BlockSize: ~128kb
The issues are more subtle beyond hardware though - your OS kernel has to understand the detection from the BIOS / EFI and if the motherboard manufacturer decided not to opt for wiring in the checksum fail signal you will run blind to data corruption that’s potentially really subtle on typical consumer hardware. With some BIOSes RAM failures result in a hard panic and will hard reboot the machine (the idea being a crash is better than continuing with an error).
For me, ECC (both UDIMM and RDIMM) today are so inexpensive it’s a no-brainer for builds for business.
"The SDRAM and DDR modules that replaced the earlier types are usually available either without error-checking or with ECC (full correction, not just parity)."
Furthermore, there is error correction at the chip level. I think a fitting analogy is like this: A single-engine airplane might seem much more dangerous than a dual engine airplane, but most single-engine airplanes have dual fuel pumps, dual ignitions sources, and additional auxiliary devices sharing the same engine body. Will it ever be as safe as a dual engine plane, no, but it's not as dangerous as having no failsafes.
Modern RAM is believed, with much justification, to be reliable, and error-detecting RAM has largely fallen out of use for non-critical applications. By the mid-1990s, most DRAM had dropped parity checking as manufacturers felt confident that it was no longer necessary.
There is no source checksum.
> transmitted file's checksum will not match the source's checksum
Most applications don't checksum their files when reading them.
> a retry will likely solve the issue
You can't retry; the only copy of the edited file is corrupt, and it's not possible to (automatically, in the general case) distinguish edits you made since loading the file out of ZFS from memory corruption that happened since loading the file out of ZFS.
2. If an error occurs on ZFS the consequences are worse, because there are very limited recovery-tools available for ZFS. The official checkdisk is "just restore from tape, it's quicker than fsck anyway". Very enterprise oriented.
Now you might be okay with the above. But the risk is greater and the potential consequences are worse. Not because of ZFS being special, but because of the way it makes use of your hardware.
Also keep in mind that a lot of people pick ZFS to reduce potential corruption so the addition of ECC is par for the course.
I mean, if the ship is on fire, you want the missile control systems to shut down after all missiles have been fired.
well not exactly true. we know that the raid5 write hole won’t tolerate random power off / memory removal etc.
as someone else said, x ray test would have been better.
I’m a little disheartened to read all of the negative comments about Btrfs in this thread as well. I’ve spent a ton of time researching Btrfs in RAID 10 for deployment on my home lab (99% Linux environment) for when it’s time to expand storage and from everything I read it seemed like it was going to be a good idea. Now I’m back to wondering if I should research ZFS again.
If you've had a negative experience, it can leave a bad taste in the mouth. People do this with food too, "I once got violently sick off chicken soup, I'll never eat it again." I wouldn't be surprised if there's an xkcd to the effect of how filesystem data loss is like food poisoning.
There is a gotcha with Btrfs raid10, it's does not really scale like a strict raid 1+0. In that traditional case, you specify drive pairs to be mirrors, sometimes with drives on different controllers so if a whole controller dies, you still have all the other mirrors on other controller and the array lives on. You just can't lose two of any mirrored pair. Btrfs raid10 is not a raid at the block level, it's done at the block group level. The only guarantee you have with any size Btrfs raid10 is the loss of one drive.
I'm even using btrfs with zstd compression on my laptop since it only has a small ssd (64 GB) and it makes it a lot more usable.
* btrfs has a few features that are Really Nice and missing in ZFS: ability of have a file system of mixed drives and adding and removing drives at will, with rebalancing. With ZFS growing a file system is painful and shrinking impossible(? still?). There has been work on it recently though.
* ZFS has a IMO MUCH cleaner design and concepts (pool & filesystems); mirrored by a much cleaner and clearer set of commands. Working with btrfs still feels like an unfinished hack. As human error is still a major concern, this is not a trivial issue.
I _have_ lost data to MD, had scary issues with BTRFS, but never had issues with ZFS in 8+ years. (The fact that FreeNAS is FreeBSD based which I'm less inclined to mess with also means that I mostly leave my appliance alone.)
You probably mean something else rather than a file system. In ZFS you do not need to grow or shrink file systems at all.
And the first alpha for RAIDZ expansion became available a week or two ago. (For going from say a 6 disk RAIDZ2 vdev, to a 7+ disk RAIDZ2 vdev). Just in case anyone decides to play with this feature, the on disk format for this feature is not stable yet, only use it on test pools.
Even back in 2009 I heard some Linux enthusiasts tell me how btrfs was going to be better than zfs!
What's sad is that it should have been; the CDDL situation is really unfortunate. Honestly, even if BTRFS performance were worse, it would be worth it in order to have a fully-supported mainlined FS... but instead its reputation is for data loss, so it's dead (yes, I know it works if you're careful, but that's a terrible quality in a filesystem).
Good to know.
But also by experience, bcachefs is incredibly stable. The only issues I had was when mismatching the tools and the kernel, leading to fsck being confused when outdated compared to the kernel. But even with that, I haven never lost any data or had it even as much as hiccup.
1: Btrfs refused to mount at one point due to a bug; the helpful folks on #btrfs walked me through the process of downgrading my linux kernel to get it into a working stat eagina. At this point I switched away from btrfs.
It's also totally worth tons of memory when you use that feature with intent. If you use dedup in combination with automated snapshots you get the most space efficient, fast and reliable incremental backup solution in existence - yes it will consume your whole server, that's the cost (works best separately as a backup server).
I have ZFS on another box that failed badly but lost no data... That was with bad ram and a motherboard that was on the "do not use" list for making ZFS NAS boxes. It always recovered the errors
Currently using it in RAID mode to hold large data sets and CCTV footage for my homelab on three drives that have smart warnings for age without any issues at all for the past two and a half years and two Ubuntu upgrades
Care to link to this list. Just want to check that my board isn't on said list!
I used to hit out of space errors on rebalancing at high usage levels, but since around kernel 4.0 the only time I hit one was when I altered data duplication settings.
And ...zfs for secure workloads.
After catastrophic failure number two I decided to never run BTRFS again.
That's an anecdote sure but that's enough to never use it ever again for me.
In my opinion BTRFS is rotted at the core, I'm more interested by Bcachefs future.
One year ago I built a RAID56 with 5 used 2TB drives from eBay. Risky move, but it went smoothly so far. It's only for home storage and important stuff is backuped off site anyway. One Seagate drive died with lots a bad sectors. The replace command took really long even with the "-r" flag (don't read from replaced drive, in theory), so I ended up unplugging the drive and rebalancing from there.
I have high hopes for bcachefs. We have a real need for a modern FS with tiered caches. I backed the project but I don't have the skills or time to help.
Rendering my laptop and the running system inoperable without a hard reset.
I’ve had no such issues with ZFS, not even on the exact same hardware and Ubuntu-release.
Needless to say, I’m not using btrfs anymore.
Edit: for perspective, I’ve never had this btrfs issue on desktops or servers. Laptops only. Might be suspend/resume related?
I've had one failure with ZFS that required developer help (space map corruption) but I got all my data back.
I wouldn't trust it's RAID features though.
Initially I had performance issues due to the default block size back then (4 KiB rather than 16 KiB; ended up just rebuilding the filesystem). There were some other issues back then regarding rebalancing and scrubbing stability, and sometimes I would have to run "btrfs-zero-log" before mounting, but I haven't had those sorts of issues for a while.
I've had multiple drive failures on my systems, and btrfs seemed to handle them as expected. I use "raid1" for metadata and "single" for data. As far as I can tell, all of the errors were due to bad blocks on physically failing drives, and btrfs was able to indicate which files were affected in all of those cases.
I've also used it to "fix" the Raspberry Pi SD card corruption issue by just running btrfs in "dup" mode—prior to that, the SD card would randomly end up with blocks being zeroed and the system would obviously start to crash, whereas btrfs just fixes the blocks up in place as it accesses the data the next time.
The NAS had two WD Red drives fail over the years with bad blocks. Not at the same time. They were detected and I replaced them with new bigger drives. Since 2012 I added more drives until now it is at 6 drives in RAID10: 4x6 TB and 2x4 TB.
I've had out of space errors on my laptops which I had to repair by adding more storage so I could successfully rebalance. Stealing the swap partition worked great for that.
Never lost any data or had any corrupt files.
Also, I always build with quality Gold standard PSUs, ECC RAM, and run my systems always plugged into an APC UPS. So I've never tried to recover a system that crashed after a lightning storm, built with WD Green drives in external USB enclosures plugged into a $2 power-strip and an HP "desktop" built out of an old laptop motherboard and the cheapest PSU HP could dig out of the trash pile.
And then in another instance it got all confused and impossible to mount read-write. It kept trying to resume some kind of transaction that wouldn't complete even with dozens of hours of CPU time.
But on the plus side it deduplicates properly, without extreme overhead.
No problems in the past few years, though it's only seeing a few hours uptime per week. It just acts as backup for data stored elsewhere, so a crash would be annoying, but not fatal - but I'm contemplating getting some new 4TB disks (now that I have a suitable controller) and putting a Plex on the thing.
This was like a year ago and I still haven't cleaned it all up due to: (1) lack of a third drive to restore stuff to (I'd prefer to leave the damaged filesystems in read-only), (2) a lack of time, (3) there now seems to be a bug in the kernel drive for my particular drives (perhaps it wasn't a power outage after all?), and (4) because I now live far away from the physical location of the drives.
Instead of using Btrfs to fix problems, I am now looking for the simplest possible solution. I was thinking of just using plain old ext4 and relying on one local and one off-site backup. It will be more work to manually look at the state of my files when a failure happens, but with something like Restic I'm at least confident that any completed backups are sound (as well as secure: Restic is the first system that I've found to be efficient while also trusting the crypto enough to back up to untrusted locations). The only open question is what to do about bit rot on the main system, since any bit rot on ext4 would just be backed up as if it was the original data. So then... maybe I'll go with Btrfs after all, using only the checksumming feature should not cause any bugs right? I haven't decided yet. First up is trying to fix this driver issue somewhere next week.
It is curious the performance differences found between ZFS and Btrfs, as I've always had the reverse experience, with ZFS being slower by maybe 15%. The scrubbing on md raid does take a while, every block must be checked as it has no idea what blocks are in use or not; although a write-intent bitmap would avoid a complete resync after an unclean shutdown.
This form of snapshotting also scales up well. I've had hundreds of snapshots with no performance reduction, and deleting them is fast also. In cases with many changes in between snapshots, it results in much more complicated metadata. And now while making a snapshot is still cheap and fast, deleting older ones starts to become more expensive for the backref walk.
The reason for this actually has nothing to do with data protection (although not having to read your entire disc set to rebuild is nice). The reason is that it's hard to figure out who RAID5/6 are actually for. The enterprise is all on RAID10 (hence no one fixing the BTRFS write hole). So you'd think it would be for enthusiasts who don't want to purchase as many discs, right? Well, in my case at least, parity disc modes are useless for me because it means I would have to buy discs of exactly the same size into the indefinite future!
I started with 4TB discs, then I've added 8TB discs, and now I'm looking at newer Western Digital 10TB helium discs, which are actually much lower power than the 8TB ones. But none of those 4TB or 8TB discs have failed! So while in theory if stick to the drive size you start out with, you can save money with RAIDZ by using parity discs and adding a drive when you need more storage (at the cost of a decreased parity %), practically speaking a lot of that money is wasted since you either have to neglect larger more efficient drives, or replace working older drives when you want to upgrade.