ZFS 2.2.1: Block Cloning disabled due to data corruption
github.com
github.com
Holding off on 2.2 seems recommended, and if you're keeping critical data on OpenZFS it might be a good idea to give the issues a glance. The 2nd one might have the same underlying solution as an issue that has given me system freezes when closing in on ~90% pool usage using 2.1 on top of LUKS.
- https://github.com/openzfs/zfs/issues/15140 : supposed to be from flushing largepages
- https://github.com/openzfs/zfs/issues/15275 fixed by https://github.com/openzfs/zfs/issues/15464 : linked to block cloning
Block cloning is a large suspect.
That said, for my main workstation, I plan to migrate to bcachefs very quickly once it's mainlined. I haven't done enough introspection to be able to tell you why I can hold both opinions in my head so well.
Another issue I ran into is when SQLite is run in synchronous mode (the default) with WAL (not default but recommended) and the default locking mode (so it creates an -shm file) and the shm file is stored on a ZFS filesystem. SQLite frequently executes ftruncate on the shm file when different processes access the same SQLite database, and for some reason ZFS can cause the ftruncate call to block until the txg timeout which is usually 5 seconds [0]. If you were running a program which records every shell command you run to a SQLite database, for example, that would cause a 5 second hang before any shell command would execute [1]. The workaround is to disable synchronous in SQLite or the ZFS dataset, which is probably safe because of the ZFS atomicity guarantees.
Those are two examples I have run into recently. I'm sure people run into issues with e.g. ext4 as well but I think they are a bit more frequent with ZFS on Linux, especially if you use more fringe features like encryption.
That said, most of the developers are likely not using it on their own systems, which probably allowed bugs to stick around abnormally long compared to the rest of the code base (although some really ancient bugs were found last year in other areas).
Corresponding HN discussion at the time: https://news.ycombinator.com/item?id=16797644
But yes I plan to use it soon-ish as well.
That said, I think he is laying a good foundation (from what I have heard/read about his development process), but there are many times when the best of us have been confident in code that turned out to have problems and I doubt he is an exception to this.
Could you please share some blog or resource to read a bit about his dev process myself? That stuff usually interests me, to improve my own process.
Anyway, it has a while since I read anything about bcachefs, but what I have read stuck me as being consistent with doing things well. For example, he is working on having an automated test suite in place before he ships it, which is a great thing to see:
For mixed storages, bcachefs will be interesting.
For RAID1, another option is any filesystem over dm-integrity over mdadm: dm-integrity can protect against silent file corruption if used below the mdadm level: any filesystem reading inconsistent data + checksum at dm-integrity level would cause dm-integrity to give a EILSEQ to mdadm, which should recover data from the mirrors.
It's done with a cryptsetup step, and explained on https://gist.github.com/MawKKe/caa2bbf7edcc072129d73b61ae781...
Main advantages:
1. it's a "bring your own filesystem" (ex: XFS, EXT4): you add protection against bitrot to any filesystem
2. dm-integrity may be newer, but mdadm and XFS have a large user base, making them well tested.
3. you can test this approach by simulating bitflips on the underlying data device, reading from the /dev/mapper entry, reading files themselves, doing a scrub etc.
4. you can select other algorithms besides crc32: in the rare case the error couldn't be seen by crc32 (which is likely to be applied at the hardware level) you gain an extra layer of safety
ZFS just is, no need to worry about layering device mappers and LVM. The automount is also very nice.
I do that, and I've even personally experienced a few rare ZFS bugs that seem due to interactions between ZFS and Western Digital firmwares.
Still, I was caught unprepared: I have several backups not stored on ZFS, but all of them were made FROM a ZFS source, meaning they are now all suspicious since silent corruption has been possible since version 2.1.4, and maybe even longer than that.
ZFS is practical to use, but for now I think I'll keep a history of file checksums, like how it was done before bitrot protection.
I'm also assuming those backups aren't actually ZFS streams (from zfs send|receive) which is a special case of "bugs biting you twice" :P
It could be a subjective feeling, but the recent years of OpenZFS development have reminded me a bit of the OpenStack experience. It seemed like almost anyone could contribute, sometimes resulting in features that were questionable in terms of stability, development, and overall thoughtfulness. Perhaps this is why iXsystems has taken a more cautious (albeit slower) approach in enabling new features in TrueNAS.
I've been running my current 32TB ZFS storage array on an HP DL380 in my basement. The power consumption and noise are both incredible.
I'm closley watching black friday/monday deals hoping for something on the Synology DS1522+.
This hit Linux first, which is why Linux experienced the issue. FreeBSD does not currently have reflink cp.
Edit: yep, it's an older bug. Set zfs_dmu_offset_next_sync=0 to save your data
https://github.com/openzfs/zfs/issues/15526#issuecomment-182...
https://github.com/openzfs/zfs/issues/15526#issuecomment-182...
https://github.com/openzfs/zfs/pull/15571/files?diff=split&w...
https://github.com/freebsd/freebsd-src/commit/068913e4ba3dd9...
Because it was introduced by a FreeBSD developer, and FreeBSD's fs and cp are a generation behind Linux's so they weren't even able to hit this.
>I don't think this ever would have happened pre-OpenZFS, and undermines the stable reputation of ZFS.
Of course it happened pre-OpenZFS: https://blog.lastinfirstout.net/2010/04/bit-by-bug-data-loss...
From the people that brought you Slowlaris.
>OpenZFS needs to do better and look to FreeBSD developers, not Linux developers, as role models.
This was introduced by a FreeBSD developer and was caught by a Linux one. You need to rethink your role models.
This is often regarded as a good thing in the BSD world. In all fairness, this bug just enforces that notion.
And Solaris (at least on amd64) wasn't slow. It was demanding, and not really suited for desktop. And it was picky with hardware.
Totally agree on the rest of your comment, though.
Same in Linux, which is why you'd choose a distro that's less bleeding edge.
>And Solaris (at least on amd64) wasn't slow.
No, it was slow. There's a reason why everyone in HPC/HFT/etc. moved off Solaris to Linux in the 2000s. Linux was regularly beating Slowlaris in practically every category at the end.
Honestly, this feature could have used more testing on both platforms.
Also, FreeBSD's VFS in theoretically able to support reflinks across datasets while Linux's VFS disallows that. It is not really behind.
I don't believe so, you're still not able to hardlink across file systems in BSD.
>It is not really behind.
It is behind: https://github.com/openzfs/zfs/issues/405
That said, I suspect you did not read the replies, since the code in question was written by a FreeBSD developer. Bugs in new code have been introduced by developers on both platforms. Unfortunately, the bugs in this feature were not caught before they reached a stable release. :/
I just learned the recent FreeBSD 14 release doesn't have 802.11n WiFi support. Is that considered bike shedding?
I have multiple zpools and none of them have been upgraded to enable any of these new features in the past couple years.