OpenZFS 2.2: Block Cloning, Linux Containers, BLAKE3
github.com
github.com
However, based on [2], it seems the checksum doesn't store the type of checksum, so it seems that it's an all-or-nothing thing instead of a setting that only applies to new blocks.
Can I actually change the checksum? What happens when I do?
Edit: It appears the type of checksum actually is stored per-block, in `blk_prop` ([3]) of `blkptr_t` ([4]).
Edit 2: The manpage says that the change only applies to new data, so yes, it's safe.
[1]: https://openzfs.github.io/openzfs-docs/Basic%20Concepts/Chec...
[2]: https://people.freebsd.org/~gibbs/zfs_doxygenation/html/d9/d...
[3]: https://github.com/openzfs/zfs/blob/c0e58995e33479a9c1d97fb2...
[4]: https://github.com/openzfs/zfs/blob/c0e58995e33479a9c1d97fb2...
One question if anyone who follows it more closely knows: are there any efforts to work specifically on ZVOL performance? They're very handy for iSCSI and other features, but my understanding is they haven't gotten much focus for quite awhile now. Maybe that just doesn't have any real company backing right now for R&D but I hope it gets some attention eventually.
I'm still patiently holding out for the first (current, usable) filesystem that let's us set a parity fraction and let me use FEC to heal small errors.
Occurs to me know that as a practical matter perhaps raw storage will solve this for a lot of people. The pace of increase in storage/$ still seems to be steady, and in turn continues to mean more and more people trivially have their total needs exceeded. At some point may be perfectly reasonable to throw half at redundancy?
If you want repair, you need to either set more than one copy in ZFS (still pretty dumb) or use multiple drives in a raidz (or mirror, though raidz is more efficient).
When you try to partition a drive like this (or configure ZFS to write multiple copies) you loose everything in the event that the whole drive fails. Honestly, it's a far better plan to just use multiple drives in a raidz so that whole-drive failures can be reconstructed in addition to data corruption.
* https://docs.ceph.com/en/latest/rados/operations/erasure-cod...
Setting copies=2 for the root filesystem, but segmenting stuff like caches, user downloads, etc into separate datasets with copies=1, is totally reasonable for this purpose. The base system in most distributions only takes a few GiB. 1 TiB SSDs are pretty common even on entry level laptops these days, and it's becoming increasingly difficult to find new machines without at least 512 GiB.
In a raidz with at least one parity disk (and any number of data disks), corruption found on any data disk can be repaired immediately. For non-overlapping corruption a single parity disk is enough. If you have corruption of the same block on two different disks (extremely unlikely) then this can be repaired if you have two parity disks. (Etc. for three identically-corrupt data disks and three parity disks)
You can have eg. two parity disks and eight data disks (or more) and so long as no more than two disks are corrupted in the same block at the same time then a scrub will repair it completely.
The system of parity disks in ZFS' raidz is meant for reconstructing entire failed data disks, so fixing corruption is no big deal.
There's this[1] work which integrated ZFS better with the kernel so it could merge small IOs on zvols better. It was merged almost a year ago. However due to a recent issue[2] it has been disabled until they can figure out what exactly is going on. If they can get it working it seemed to give a decent boost for certain scenarios.
* https://openzfs.github.io/openzfs-docs/man/master/8/zfs-clon...
* https://openzfs.github.io/openzfs-docs/man/master/7/zfsconce...
(ZFS uses the fairly 'standard' nomenclature of "snapshots" being read-only copies and "clones" being read-write copies.)
> Corrective "zfs receive" (#9372[0]) - A new type of zfs receive which can be used to heal corrupted data in filesystems, snapshots, and clones when a replica of the data already exists in the form of a backup send stream.
Does this mean I can use the overlay2 driver with docker instead of the ZFS one? I've had many people tell me this was possible in the past when it definitely wasn't.
The performance and reliability of the ZFS driver leaves lots to be desired. The last time I ran Docker on a ZFS partition I actually used a zvol formatted ext4 just to avoid these issues.
No wonder: it calls the zfs(1) tool and parses its output to do its work. Again. and again. and again.
Thing is, https://github.com/moby/moby/blob/670bc0a46c4ca03b75f1e72f73... is using https://github.com/mistifyio/go-zfs which features code like `out, err := zfsOutput("get", "-H", key, d.Name)` (Source: https://github.com/mistifyio/go-zfs/blob/master/zfs.go#L315) to get a single zfs property.
Somebody chose to use a library as abstraction that looks good but is implemented as a MVP (nothing wrong with that). "In the future, we hope to work directly with libzfs" should have raised an alarm somewhere, though.
Looks like the answer to that is "yes, but be careful as some are kind of dodgy". :/
---
Oh, that library bills itself as "Simple wrappers for ZFS command line tools", rather than as a ZFS interface.
Wonder why the Moby project picked that one then? Maybe a case of "it was the best choice at the time".
ZFS 2.2.0 (RC): Block Cloning merged - https://news.ycombinator.com/item?id=36588240 - July 2023 (165 comments)
Can drives be removed in the same way if there's the space for it?
It is your hardware where the unreliability lies and ZFS detects that. Would you rather prefer silent corruption?
Otherwise the feature set seems great: both a volume manager and a file system, seamless compression, deduplication, COW and snapshots - what's not to love :)
But there is. It's `zpool clear <poolname>`.
That said, you probably need to have created or imported the pool using stable device names (e.g. `zpool import -d /dev/disk/by-id` or `by-part-uuid`, for example). Otherwise, when reattaching the USB device it might get assigned a different device name (e.g. `/dev/sdb` instead of `/dev/sda`) and ZFS might think the device is still unavailable.
Okay
If the drive lies and says that part of the journal has been written when it hasn't yet, and ZFS goes ahead and writes the next part of the journal, then when you unplug the drive and the first part of the journal goes away (which the 2nd depended on) you're hosed. There have to be places where ZFS blocks until something critical has definitely, absolutely, been written to disk.
At that point the only thing ZFS can do is try to unwind back to whatever it thinks is a consistent state, but this isn't 100% guaranteed. (it depends on what old data is still hanging around)
That said, I have you tried something like this[1]? Also what device name did you use when creating the pool? Using one of the /dev/disk/by-'s that doesn't change when you reconnect would make it a lot smoother I imagine.
[1]: https://github.com/openzfsonosx/zfs/issues/104#issuecomment-...
Like suppose you're copying a large genome from A to B, and you already have a genome from that species (but different organism) lying around on B.
Sure, you could clone and then rsync on top of the clone to avoid transferring the common bits again. But that's forethought that users often don't have. Better to use the rolling hash related metadata that rsync generates as a query into all possible targets and pick one automatically so that the user doesn't have to think about it and just sees it as a really fast copy.
This would be automatic. If there's similarity on the target, use it, but without requiring the user to tell you where to find it.
While block cloning isn't supported for encrypted datasets yet, looks like there is a WIP already: https://github.com/openzfs/zfs/pull/14705
Very cool to have the ability for near instant cp or mv between datasets or snapshot (at least locally; the links don't persist with send/recv?).
I wonder what holds back a ZFS-level offline dedupe function now that that's implemented since you could already basically write a shell script to do something like it.
https://learn.microsoft.com/en-us/windows-server/storage/ref...
[1] https://www.reddit.com/r/DataHoarder/comments/iow60w/testing...
I don't think folks should default to using one or the other, everything depends on your workloads. Ext4 might even be the best choice. Though bcachefs could be an "endgame FS" if it delivers on its promises.
This article is somewhat outdated but the author is knowledgeable on the subject https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
ZFS RAID is production quality and safe to use. Not the case with Btrfs:
* https://btrfs.readthedocs.io/en/latest/btrfs-man5.html#raid5...
Certainly striping-mirroring will give you more IOps, but if you want more space efficiency for bulk storage, RAID-y solutions is probably better.
ZFS has been in production use since Solaris 10u2 (June 2006), so we're approaching twenty years of it being banged on.
Btrfs works and has worked fine for RAID0 and RAID1. It is specifically RAID5, RAID6, et al.
That said, I prefer ZFS (we'll see what happens when bcachefs is merged), though Synology DSM doesn't offer it and it can be a PITA to not being able to run latest Linux kernel. Especially (solely?) on a rolling distro I didn't like that.
There's drivers for both to run under Windows btw.
I have to say as a Linux user I find the ZFS tooling surprisingly difficult, and I've had far more issues with Linux on ZFS root than BTRFS. YMMV. On my one BSD machine I currently have a usb-attached zpool that seems completely frozen in spite of reboots for over a week -- all zpool and zfs commands hang indefinitely with no output (even though the machine is otherwise working fine). No idea what to make of that, was working fine for 3+ years previously.
That said, people far more knowledgeable than me far prefer ZFS, so I keep my really important long-term storage and backups on a RAIDZ2 array.