Maintaining sufficient free space with ZFS
taras.glek.net
taras.glek.net
But, of course, there is no defrag for ZFS and so we do it the trailer-park way - we add free space to a zpool when adding a vdev, and then we 'zfs send' datasets to ourselves on that same zpool thus allowing ZFS to lay down those bytes in an efficient and orderly fashion ... as opposed to the inefficient way they were laid down over time.
This works.
https://rsync.net/resources/notes/
... we've not disussed this item yet but I am finalizing the most recent "notes" and will include it.
If you don't have busy filesystems filling beyond ~90% (used to be 80%...) then I wouldn't ever worry about this.
OTOH, if you have very busy filesystems with lots of different consumers and hundreds of millions (or billions) of inodes then a recent dataset might be scattered all over your zpool. There are absolutely performance ramifications to this.
So you expand that zpool (you have to expand it for this to work) by adding a vdev, and then you 'zfs send' that dataset onto the same zpool and it will be laid down nice and orderly onto the new free space that was added to the zpool.
That dataset will then be more performant. This is analogous to "defrag" as we think of it but, again, in a duct-taped-engineering kind of way.
Why not? Seems like an important thing to have.
The fundamental problem is that ZFS is not content addressed. It's almost CAS, but not quite. In ZFS block pointers have physical addresses in the pointer. That means that physical block locations are inexorably part of the ZFS Merkle hash tree. And that means that any change to the location of any block necessitates that the pointers to it must change and so be rewritten, but now you have to find all those pointers -- even if there can be only one pointer, you still have to change it, and when you change it, you change the Merkle hash tree of the dataset.
The solution, IMO, is to split all znodes and any interior nodes that have block pointers into two halves. One half should have only logical block pointers free of any physical location pointers, thus the pointers in this half should be some metadata + hashes of the pointed-to blocks. The second half should be a "cache" of physical locations corresponding to the logical block pointers in the first half, plus a hash of just those physical locations. The key is that the Merkle hash tree should not bind physical locations so that changing those locations does not alter the Merkle hash tree. That leaves only the task of updating those second block halves that carry physical locations. So when traversing the tree the filesystem could optimistically read pointed-to blocks from the cached locations, and if the hash of the block read does not match the logical pointer, then go looking in a log of recent block moves. This way a BP rewrite system could simply traverse a dataset looking for blocks to move, copy the blocks to new locations, update the cached locations where the block pointer was found, log the move, then add the old location to the freelist.
But ZFS is not like that. ZFS is not a content-addressed store, but it's a Merkle hash tree all the same. That makes BP rewrite insanely too difficult.
I do wonder how much effort it would take to make a ZZFS that is based on ZFS but with the CAS redesign sketched above. This is not the first time I've written about this. I think I first proposed this back in the early days of Illumos, and one or two people thought that not hashing the physical locations into the Merkle hash tree would be folly -- I think they're wrong, because either you trust your hash function or you don't. Granted, to make ZFS into a proper CAS does require a strong cryptographic hash function, but ZFS uses one so...
... and this was much more useful and interesting :)
Instead I think the OpenZFS community could maybe develop an automatic and transparent feature like what you do: internally "zfs send" a dataset, and then apply any transactions that took place at the origin while doing that, then lock the origin dataset, do one more round of transaction copying to the new dataset, then atomically swap the old and new, unlock, then destroy the old.
But I'm very curious what filesystems devs think of the CAS idea. It's not original, mind you. CAS is a well researched topic. I'm dead certain there's nothing wrong with a CAS design for a filesystem -- the key is to make the logical -> physical block pointer mapping fast, which is why I'd have a cache right next to every interior block, with the cache excluded from the Merkle hash tree.
I would be curious to see how practical someone writing something like this would be without requiring one make an entirely new pool or do offline conversion...
The fundamental problem I described above is just too hard to overcome in ZFS. Therefore the only thing that can be done about problems like this is to find alternative solutions. vdev evac is one such solution to one of the problems BP rewrite was meant to solve. In a reply to a sibling to yours I outline a way in which ZFS could solve the defragmentation problem without a proper BP rewrite. Enough such solutions might make BP rewrite not necessary at all.
BP rewrite is still needed for defragmentation, and for other reasons too (like if you wanted to change the configuration of a zpool to have more mirrors or stripes, or if you wanted to change compression algorithms, etc).
(A proper CAS would not be able to handle things like compression algorithms, unfortunately, not unless the filesystem hashed the decompressed block rather than the compressed block, but that has security issues if you don't trust the storage, so it's not ideal. Of course, if you don't trust the storage then you need to be signing/MACing Merkle hash tree roots, but that's another story.)
The problem with this change is not that it's breaking, but that it essentially forks too many code paths in ZFS to be worth doing.
I'm not aware of a great write-up of this, but btrfs uses a "chunk tree" to map extents of logical addresses to >=1 physical stripes.
https://github.com/btrfs/btrfs-dev-docs/blob/master/tree-ite...
But it's true that it has some bad performance properties. The ZIL is essentially a way to amortize what would otherwise be very expensive b-tree transactions -- expensive because every interior node on the path to the block you're trying to write also needs a new write, so you get O(depth) write magnification, which means write performance becomes 1/depthth of storage write performance, which is awful. But the ZIL properly amortizes all those interior node writes, making it possible to do just one write of each of those for any number of leaf node writes that fit in the space of time between full transactions.
So a ZIL-like log is essential and makes CAS write performance tolerable.
My dream is to be able to use a TPM to hold a key for the whole zpool that can't be recovered unless you boot into a blessed dataset snapshot whose root hash is part of the TPM key unlock policy, then combined with other bits of secure boot technology and remote attestation (this latter for enterprises, not individuals) you'd get a pretty secure-against-physical-theft setup.
A list of 'missing' features that need data moving off the top of my head: defrag (since we're talking about it), shrinking partitions (has some overlap with defrag), compression of existing data, dedup of existing data, moving data between filesystems on the same pool without having to read the data and write it a second time. I think there's maybe one other gotcha I can't remember.
All these things might be nice, but are tricky to get right, and users can figure out other ways to do them, so there you go. Limited resources and all.
For example, we've had three different zlibs compressing gzip streams for a while, and nobody really complained about it. (Linux ships 1.1.x; FBSD and everyone else ships 1.2.x; Intel QAT produces identical output to neither.)
For a while, zstd was outputting different results on BE and LE systems and nobody noticed.
Most of the reason we haven't updated compressors nowadays is that nobody's convinced anyone it's got enough benefit to be worth the added complexity.
At which point fragmentation usually isn't that big of a deal. But sure, if you add enough new storage it works.
It's more of a problem for individuals. That and that you can't increase the size of a vdev is (in the works but with lots of caveats) really dampens my enthusiasm as a hobbyist.
It's not that ZFS takes a long time to update free space - it's that it actually doesn't necessarily free the space immediately.
Deletes over a certain size (by default 20480 blocks, so 10/80 MB for 512b/4k sectors, respectively, I believe) dump the thing to be deleted (assuming it's to be freed, dedup refcounting/snapshots/etc might mean it's not actually freed) on an async queue that gets worked through in the background, rather than blocking rm on it.
(Otherwise, if the files are smaller than that, I think it's as the author says, with bundling things into a txg and only updating once they're flushed out. But the example was ~40 files totalling several GB, so I think it's as I describe above.)
A process to delete less important files to maintain free space that waited on the delete queue to drain before reassessing the situation would need something more sophisticated.
Could be wrong. Maybe I'll remember and go try it in a day or two, when I'm not at an event[1]...
[1] - https://openzfs.org/wiki/OpenZFS_Developer_Summit_2022
The first time you delete something bit and it doesn't show up in df right away, it might be surprising though. Or this case where automation deleted more than expected because the author wasn't aware of the need to wait.
If system is busy, sure, batching requests helps performance, especially on spinning rust.
In one that it isn't ? Why wait ? Then again checking space after every removed file (which I assume what the problem was here) instead of "generate list of files to remove, remove them, then go to sleep for an hour" is a bit suboptimal in the first place
Using --track-bytes-deleted on ZFS wont really track the bytes deleted, but just the size of the file that was deleted -- as the author points out, that may not be the number of bytes actually freed.
For my money, "sleep 10 before checking free-space" is a more robust solution.
My choice would be a cron job to run every few hours (or every day) that calculates space required and deletes logs files to that size. As long as the desired free space leaves an adequate margin, this would be robust and work in the presence of compressed files, snapshots, etc.
"Sleep 10" should be fine, but it's also a hack, and the author would need to consider how long it takes for cleared space to become available on a heavily loaded system. I would expect that under severe load, the clear-up lag could be very long.
And running more often. 50GB(?) at a time can back up any filesystem.
Not to mention, there is another option which perhaps better matches this use case.
And author seems to be using the zfs quota property when the refquota property might better account for descendants and perhaps make a better initial calculation of what free space is available easier?
ZFS is great but absolutely has tradeoffs. It makes really good choices. Maybe this is one?
A very common thing to do with ZFS is to have automatically rotating snapshots, so that you can go and retrieve a file that you accidentally deleted or changed, without having to go to the backups. In this case, the space is freed when the last snapshot containing those files is rotated away, which could be day or a week later, depending on how it is set up.
Such code would break on zfs.
Is there even a better way to achieve "I want all spare space to be filled with logfiles, but I never want to run out of space."?
It would be especially good for caches of remote files or data that could be recalculated on demand.
We already have lots of similar things for RAM - for example Javas SoftReference.
A bit more turns up `zdb -e mainpool -ul` (or maybe without the -e?) that includes a bunch of txg ids along with tons of other stuff. I'd think there ought to be a less noisy way to find that, but I don't know where to start looking.
It seems to that #3 at the bottom, snapshots holding onto a file, is the biggest or most common culprit of free space not getting freed on delete. Having a good snapshot management strategy is critical if you use the feature at all. And if you're using it lazily, just know that you should delete old snapshots when going through a free-space exercise.
I can't recall if the threshold is 80% or 90% right now, but around there.
[1] - https://github.com/openzfs/zfs/blob/eaaed26ffb3a14c0c98ce1e6...
Yes, you can run busy zpools with many different random access consumers at 90%.
I wouldn't do it without a SLOG but it is really only at 92%, these days, that we start to see performance really degrade.
YMMV.
I'd be curious if you're suggesting you're running something based on 2.1 and still finding you can run it fuller more reasonably.
[1] - https://github.com/openzfs/zfs/commit/aa755b35493a2d27dbbc59...
i.e. do an unlink a 1TB file that is in use in linux on ext4. file disappears immediately form namespace but space isn't freed up. close program keeping that file opened and space will slowly be reclaimed. similarly, due to unlink / ext4 semantics, if one did unlink that file without anyone keeping it open, the unlink() would block until the file was deleted, but I'm not sure that's really required by posix (i.e. the file could be opened by someone). Therefore I think one has to believe that the expectation on any file system is that unlink returning doesn't mean the space was freed up yet.