ZFS 2.2.0 (RC): Block Cloning merged
github.com
github.com
> It is in FreeBSD main branch now, but disabled by default just to be safe till after 14.0 released, where it will be included. Can be enabled with loader tunable there.
> more code is needed on the ZFS side for Linux integration. A few people are looking at it AFAIK.
- Linux container support (So overlayfs is ZFS aware)
- BLAKE3 checksums (So faster checksums)
- Zstd early abort (I though Zstd already had early abort! This means you can enable Zstd compression by default, and if the data isn't compressible, it aborts the operation and stores it uncompressed, avoiding the extra CPU use.
- Fully adaptive ARC eviction (Separate adaptive ARC for data and for metadata. Until now metadata ARC was almost FIFO and not efficient)
- Prefetch improvements (That's always nice)This was already implemented for all compression algorithms since ZFS was released as open source in 2005, at least (and even though zstd didn't yet exist then, it still benefited from this feature when it was integrated into ZFS).
The feature you describe only saves CPU time on decompression, but still requires ZFS to try to compress all data before storing it on disk.
This new early abort feature actually seems to be even more clever/sophisticated, because it also saves a lot of CPU time when compressing, not just descompressing.
Usually you are forced to select a fast compression level so that your performance doesn't degrade too much for little benefit, because if you had a higher compression level, ZFS would always be paying a high CPU cost for trying to compress data which often is not even compressible (since a lot of data on a typical filesystem belongs to large files, which are often not compressible).
However, with early abort, you can select a higher compression level (zstd-3 or higher) without paying a penalty for spending a lot of time trying to compress uncompressible data.
The way it seems to work is that ZFS first tries to compress your data with LZ4, which is an extremely fast compression algorithm. If LZ4 succeeds in compressing the data, then ZFS discards the LZ4-compressed data and recompresses it with your chosen higher compression level, knowing that it will be worth it.
If LZ4 can't compress the data, then ZFS doesn't even try to compress with the slower algorithm, it just stores the data uncompressed.
This almost completely avoids paying the CPU penalty for trying to compress data that cannot be compressed, and only results in about 9% less compression savings than if you had simply chosen the higher compression level without early abort.
The actual algorithm is even more clever: in-between the LZ4 heuristic and a zstd-3 or higher compression level, there is an intermediate zstd-1 heuristic that is also done to see whether the data is compressible.
This intermediate heuristic seems to almost completely eliminate the 9% compression savings cost while still being much faster than trying to compress all data with zstd-3 or higher, since zstd-1 is also very fast.
In the end, you get almost all the benefits of a higher compression level without paying almost any cost for wasting CPU cycles trying to compress uncompressible data!
To add something amusing but constructive - it turns out that one pass of Brotli is even better than using LZ4+zstd-1 for predicting this, but integrating a compression algorithm _just_ as an early pass filter seemed like overkill. Heh.
It's actually even funnier than that.
It tries LZ4, and if that fails, then it tries zstd-1, and if that fails, then it doesn't try your higher level. If either of the first two succeed, it just jumps to the thing you actually requested.
Because it turns out, just using LZ4 is insanely faster than using zstd-1 is insanely faster than any of the higher levels, and you burn a _lot_ of CPU time on incompressible data if you're wrong; using zstd-1 as a second pass helps you avoid the false positive rate from LZ4 having different compression characteristics sometimes.
Source: I wrote the feature.
I made some minor edits within 15-30 mins of posting the comment, but the LZ4+zstd-1 heuristic was there from the beginning (I think) :)
Since you wrote this feature, I've been wondering about something and I was hoping you could give some insights:
Is there a good reason why this early abort feature wasn't implemented in zstd itself rather than ZFS?
It seems like this feature could be useful even for non-filesystem use cases, right? It could make compression at higher levels substantially faster in some cases (e.g. if a significant part of your data is not compressible).
Having this feature also available as an option in the zstd CLI tool could be quite useful, I think. I was thinking it could be implemented in the core zstd code and then both the CLI tool and ZFS could just enable/disable it as needed?
Thank you for your great work!
Edit: I just now realized, based on your Brotli comment, that since these early passes can be done (in theory) with arbitrary algorithms, perhaps it's better implemented as an additional layer on top of them rather than implementing it in each and every compression algorithm.
So I just tried using LZ4 as a first pass, and that worked well enough, with a high false positive rate, so zstd-1 was duct-taped onto it.
Trying to just glue an entropy calculator on is something I want to try, but I haven't, because this worked "well enough".
As far as why not do it in just zstd or LZ4 proper...well, this works very well for extremely quantized chunks, like ZFS records, but for things with the streaming mode, it becomes more exciting to generalize.
Plus it's much nicer for the ZFS codebase to keep it abstracted.
Also, as far as using it for, say, gzip...well, I should really open that PR from that branch I have, shouldn't I? ;)
I bought a super cheap QAT card with the idea of doing hardware accelerated gzip compression on ZFS (just for playing with it), but seems like the drivers are kind of a mess, and not being a developer who could debug it, maybe I should stay with the easy and fast zstd with early abort :P :) Less headaches .
I thought Intel stopped selling QAT cards with their newer generations of it?
Personally I find that gzip is more or less strictly worse than zstd for any use case I can imagine, though if you can hardware offload it, then...it's free but still worse at its job. :)
Though zstd can apparently be hardware offloaded if you have like, Sapphire Rapids-era QAT and the right wiring...sometimes. (I believe the demo code says it doesn't work on the streaming mode and only works on levels 1-12.) [1]
I think QAT API used to be a mess in past generations, and IIRC it's one of the factors drivers for older hardware weren't included in TrueNAS Scale. I think it was this ticket [1] where they said that it was complicated to add support for previous hardware, but can't confirm because JIRA's been acting funny since yesterday.
Feels like QAT has stabilized since this Gen4 hardware and interesting things are appearing again. Not only the ZSTD plugin you linked, but I just noticed QATlib added support for lz4/lz4s past July [2]! That's interesting.
They are also publishing guides for accelerating SSL operations in SSL terminators / load balancers like HAProxy [3], so between this and cloud providers using DPUs to accelerate their load, doesn't feel like they're shutting this down soon.
I'd love if they make this a bit more accessible to non-developers.
[1] : https://ixsystems.atlassian.net/browse/NAS-107334
[2] : https://github.com/intel/qatlib/tree/main#revision-history
[3] : https://networkbuilders.intel.com/solutionslibrary/intel-quickassist-technology-intel-qat-accelerating-haproxy-performance-on-4th-gen-intel-xeon-scalable-processors-technology-guideAs far as wasteful - not really? It might be possible to be more efficient (as I kept saying in my talk about it, I really swear this shouldn't work in some cases), but LZ4 is so much cheaper than zstd-1 is so much cheaper than the higher levels of zstd, that I tried making a worst-case dataset out of records that experimentally passed both "passes" but failed the 12.5% compression check later, and still, the amount of time spent on compressing was within noise levels compared to without the first two passes, because they're just that much cheaper.
I tried this on Raspberry Pis, I tried this on the single core single thread SPARC I keep under my couch, I tried this on a lot of things, and I couldn't find one where this made the performance worse by more than within the error bars across runs.
(And you can't use zstd-fast for this, it's a much worse second pass or first pass.)
Have been looking forward to this for years!
This is so much better than automatically doing dedup and the RAM overhead that entails.
Doing offline/RAM+in memory dedup size optimizations seem like a really good optimization path. In the spirit of also paying only what you use and not the rest.
Edit: What's the RAM overhead of this? Is it ~64B per 128kB deduped block or what's the magnitude of things?
No real memory impact. There's a regions table that uses 128k of memory per terabyte of total storage (and may be a bit more in the future). So for your 10 petabyte pool using deduping, you'd better have an extra gigabyte of RAM.
But erasing files can potentially be twice as expensive in IOPS, even if not deduped. They try to prevent this.
The answer to that is messy, but basically, there's a table that should be kept in memory for "is there anything reflinked at all in this logical range on disk", and that covers large spans, so how many entries would depend on how contiguous your data logically was on disk; the actual precise mapping list per-vdev doesn't need to be kept continuously in memory, just the more coarse table, so that saves you a fair bit on memory requirements.
This largely exists so you can avoid doing extra work on free.
No gotchas / issues, works well, easy to setup.
I am looking forward to the Direct IO speed improvements for NVMe drives with https://github.com/openzfs/zfs/pull/10018
edit: one thing I forgot to mention is, when creating your pool make sure to import your drives by ID (zpool import -d /dev/disk/by-id/ <poolname>) instead of name in case name assignments change somehow [1]
[1] https://superuser.com/questions/1732532/zfs-disk-drive-lette...
See also "Scaling ZFS for NVMe" by Allan Jude at EuroBSDcon 2022:
I’ve been out of this particular game for a long time.
Note that we do not use ZFS for content, since it is incompatible with efficient use of sendfile (both because there is no async handler for ZFS, so no async sendfile, and because the ARC is not integrated with the page cache, so content would require an extra memory copy to be served).
One thing that makes the interfaces in Linux much messier than the FreeBSD ones is that a lot of the core functionality you might like to leverage in Linux (workqueues, basically anything more complicated than just calling kmalloc, and of course any SIMD save/restore state, to name three examples I've stumbled over recently) are marked EXPORT_SYMBOL_GPL or just entirely not exported in newer releases, so you get to reimplement the wheel for those, whereas on FreeBSD it's trivial to just use their implementations of such things and shim them to the Solaris-ish interfaces the non-platform-specific code expects.
So that makes the Linux-specific code a lot heavier, because upstream is actively hostile.
I wish someone would come in and convince the kernel devs that "hey, if you want EXPORT_SYMBOL_GPL to have legal weight in a copyleft sense then you can't just slap it onto interfaces for political reasons"
IMO Linus should stop being half-and-half about it and either mark everything SYMBOL_GPL and see how well that goes or stop this nonsense.
Linus himself also made remarks about ZFS at one point that were pretty...hostile. [1] [2]
> The fact is, the whole point of the GPL is that you're being "paid" in terms of tit-for-tat: we give source code to you for free, but we want source code improvements back. If you don't do that but instead say "I think this is _legal_, but I'm not going to help you" you certainly don't get any help from us.
> So things that are outside the kernel tree simply do not matter to us. They get absolutely zero attention. We simply don't care. It's that simple.
> And things that don't do that "give back" have no business talking about us being assholes when we don't care about them.
> See?
Note that there's at least one unfixed Linux kernel bug that was found by OpenZFS users, reproducible without using OpenZFS in any way, reported with a patch, and ignored. [3]
So "not giving back" is a dubious claim.
[1] - https://arstechnica.com/gadgets/2020/01/linus-torvalds-zfs-s...
[2] - https://www.realworldtech.com/forum/?threadid=189711&curpost...
(/s)
Also, ZFS goes with the same licensing interpretation as Linus used with AFS.
Which three?
As far as know the point of EXPORT_SYMBOL_GPL was to push back on companies like Nvidia who wanted to exploit loopholes in the GPL. That seems to me like a reasonable objective.
Relevant Torvalds quote: https://yarchive.net/comp/linux/export_symbol_gpl.html
But if you're marking interfaces as GPL-only, or implementing taint detection that means if you use a non-SYMBOL_GPL kernel symbol which calls a GPL-only function it treats the non-SYMBOL_GPL symbol as GPL-only and blocks your linking, it gets a bit out of hand.
Building the kernel with certain kernel options makes modules like OpenZFS or OpenAFS not link because of that taint propagation - because things like the lockdep checker turn uninfringing calls into infringing ones.
Or a little while ago, there was a change which broke building on PPC because a change made a non-SYMBOL_GPL call on POWER into a SYMBOL_GPL one indirectly, and when the original author was contacted, he sent a patch reverting the changed symbol, and GregKH refused to pull it into stable, suggesting distros could carry it if they wanted to. (Of course, he had happily merged a change into -stable earlier that just implemented more aggressive GPL tainting and thereby broke things like the aforementioned...)
All ZFS needs to do is just have one of Oracles many lawyers say "CDDL is compatible with GPL". Yet, they Oracle don't.
It's explicitly not compatible with GPL, though. It has clauses that are more restrictive than GPL, and IIRC some people who contributed to the OpenZFS project did so explicitly without allowing later CDDL license revisions, which removes Oracle's ability to say CDDL-2 or whatever is GPL-compatible.
So even if someone rolled up dumptrucks of cash and convinced Oracle that everything was great, they don't have all the control needed to do that.
But "call an opaque function that saves SIMD state" is obviously not derivative of the kernel code in any way. The more exports that get badly marked this way, the more EXPORT_SYMBOL_GPL becomes indistinguishable from EXPORT_SYMBOL.
The main "legal effect" I see is that you are not willing to take that risk, just like Oracle isn't.
If enough symbols get restricted without valid copyright-based reasons, modules that legitimately have non-GPL licenses will have to lie to the kernel to get loaded. And that's a stupid situation all around.
The ZoL project lead said at one point there were a variety of reasons this wasn't initially done for the Linux integration [1], but that it was worth taking another look at since that was a decade ago now. Having looked at the Linux memory subsystems recently for various reasons, I would suspect the limiting factor is that almost all the Linux memory management functions that involve details beyond "give me X pages" are SYMBOL_GPL, so I suspect we couldn't access whatever functionality would be needed to do this.
I could be wrong, though, as I wasn't looking at the code for that specific purpose, so I might have missed functionality that would provide this.
[1] - https://github.com/openzfs/zfs/issues/10255#issuecomment-620...
But sure, it's certainly an older issue, and given that the ABD rework happened, I wouldn't put anything past being "feasible" if the benefits were great enough.
(Look at the O_DIRECT zvol rework stuff that's pending (I believe not merged) for how a more cut-through memory model could be done, though that has all the tradeoffs you might expect of skipping the abstractions ZFS uses to minimize the ability of applications to poke holes in the abstraction model and violate consistency, I believe...)
[0] https://www.kernel.org/doc/Documentation/filesystems/dax.txt
Back to the topic at hand, it’s actually scary how few software expose control over whether or not sendfile is used, assuming support is only a matter of OS and kernel version but not taking into account filesystem limitations. I ran into a terrible Samba on FreeBSD bug (shares remotely disconnected and connections reset with moderate levels of concurrent ro access from even a single client) that I ultimately tracked down to sendfile being enabled in the (default?) config - so it wasn’t just the expected “performance requirements not being met” with sendfile on ZFS but even other reliability issues (almost certainly exposing a different underlying bug, tbh). Imagine if Samba didn’t have a tubeable to set/override sendfile support, though.
LLNL openzfs project: https://computing.llnl.gov/projects/openzfs Old presentation from intel with info on what was one of our bigger deployments in 2016 (~50pb): https://www.intel.com/content/dam/www/public/us/en/documents...
I believe it was called ZFS On Linux or something like that.
Nice how things have evolved: from FreeBSD to linux and back. In my mind this has always been a very inspiring example of a public institution working for the public good.
ZoL, if my ancient memory serves, was at LLNL, not based on the FreeBSD port (if you go _very_ far back in the commit history you can see Brian rebasing against OpenSolaris revisions), but like 2 or 3 different orgs originally announced Linux ports at the same time and then all pooled together, since originally only one of the three was going to have a POSIX layer (the other two didn't need a working POSIX filesystem layer). (I'm not actually sure how much came of this collaboration, I just remember being very amused when within the span of a week or two, three different orgs announced ports, looked at each other, and went "...wait.")
Then for a while people developed on either the FreeBSD port, the illumos fork called OpenZFS, or the Linux port, but because (among other reasons) a bunch of development kept happening on the Linux port, it became the defacto upstream and got renamed "OpenZFS", and then FreeBSD more or less got a fresh port from the OpenZFS codebase that is now what it's based on.
The macOS port got a fresh sync against that codebase recently and is slowly trying to merge in, and then from there, ???
Some discussion: https://www.reddit.com/r/NixOS/comments/ops0n0/big_shoutout_...
It can be made somewhat[0] stable with using no (or a simple compression like ZLE), avoiding log devices and caches it can work, but it's way simpler and guaranteed stable to just use a separate partition.
[0] even setting the myriad of options there still exist reports about hangs, like https://github.com/openzfs/zfs/issues/7734#issuecomment-4167...
On a VM host you have a bit of leeway for memory overcommit, but ideally if you have a box that frequently has to rely on swap, something is bad wrong. I don't think it's really fair to criticize a worst-choice configuration for being poorly supported here. There are simply too many opportunities for catch-22/deadlock scenarios in the interaction of the fs and allocator during OOM...
We typically use Proxmox. It’s a convenient node host setup and usually has a very up to date zfs and it’s stable
I just wouldn’t use the Proxmox web ui for zfs configuration. It doesn’t have up to date options. Always configure zfs on the cli
What options are missing for you?
Would be great if you could open an enhancement request over at https://bugzilla.proxmox.com/ for tracking this (no promises on (immediate) implementation though).
They have both FreeBSD and Linux based stuff, targeting different use cases.
We've successfully survived numerous disk failures (a broken batch of HDD's giving all kinds of small read errors, an SSD that completely failed and disappeared, etc), and were in most cases able to replace them without a second of downtime (would have been all cases if not for disks placed in hard-to-reach places, now only a few minutes downtime to physically swap the disk).
Snapshots work perfectly as well. Systems are set up to automatically make snapshots using [1], on boot, on a timer, and right before potentially dangerous operations such as package manager commands as well. I've rolled back after botched OS updates without problems; after a reboot the machine was back in it's old state. Also rolled back a live system a few times after a broken package update, restoring the filesystem state without any issues. Easily accessing old versions of a file is an added bonus which has been helpful a few times.
Send/receive is ideal for backups. We are able to send snapshots between machines, even across different OSes, without issues. We've also moved entire pools from one OS to another without problems.
Knowing we have automatic snapshots and external backups configured also allows me to be very liberal with giving root access to inexperienced people to various (non-critical) machines, knowing that if anything breaks it will always be easy to roll back, and encouraging them to learn by experimenting a bit, to the point where we can even diff between snapshots to inspect what changed and learn from that.
Biggest gotchas so far have been on my personal Arch Linux setup, where the out-of-tree nature of ZFS has caused some issues like a incompatible kernel being installed, the ZFS module failing to compile, and my workstation subsequently being unable to boot. But even that was solved by my entire system running on ZFS: a single rollback from my bootloader [2] and all was back the way it was before.
Having good tooling set up definitely helped a lot. My monkey brain has the tendency to think "surely I got it right this time, so no need to make a snapshot before trying out X!", especially when experimenting on my own workstation. Automating snapshots using a systemd timer and hooks added to my package manager saved me a number of times.
[1]: https://github.com/psy0rz/zfs_autobackup [2]: https://zfsbootmenu.org/
I do that with sqlite to keep a selection of snapshots from the last hours, days etc.
A slightly annoying property of snapshots and clones is the inability to fully re-root a tree of snapshots, e.g. permanently split a clone from its original source and allow first-class send/receive from that clone. The snapshot which originated the clone needs to stick around forever[2]. This prevents a typical virtual machine imagine process of keeping a base image up to date over time that VMs can be cloned from when desired and eventually removing the storage used by the original base image after e.g. several OS upgrades.
I don't have any big performance requirements and most file storage is throughput based on spinning disks which can easily saturate the gigabit network.
I also use ZFS on my laptop's SSD under Ubuntu with about 1GB/s performance and no shortage of IOPS and the ability to send snapshots off to the backup system which is pretty nice. Ubuntu is going backwards on support for ZFS and native encryption uses a hacky intermediate key under LUKS, but it works.
[0] https://github.com/openzfs/zfs/issues/12014 [1] https://github.com/openzfs/zfs/issues/12594 [2]https://serverfault.com/questions/265779/split-a-zfs-clone
We also use ZFS for testing of our own block device. My biggest problems with it are related to this activity. Hot-plugging or removing block devices from the pool often leads to unrepeatable pool, and the tests have to be scrambled because the whole system needs to be rebooted.
I don't understand what you mean by this comment. Are you creating a zvol on your zpool and loopback mounting it with a different filesystem? It sounds like you are very much using ZFS as a filesystem, even if you are not booting from it or using it for any kind of high availability purpose.
I do agree with you on the matter that there should be a way to clear out an (UNAVAIL vdev / SUSPENDED pool) without a reboot. This has been a thorn in my side for as long as ZFS has existed. The deadlock-by-design line is really quite nonsense.
The PersistentVolume API is a nice way to divvy up a shared resource across different teams, and using ZFS for that gives us the snapshotting, deduplication, and compression for free. For our workloads, it benchmarked faster than XFS so it was a no-brainer.
The fact how I can replicate live data of arbitrary size (many TB sized filesystems) in small chunks to other hosts every minute greatly increased my deep sleep quality. Of course databases with multi-host write etc are nice, but in our use case, all customers are rather small with just lots and lots of files (medical and otherwise), the database itself is rather small and doesn't need replication.
Best thing, on the receiver side of the backup ZFS ensures due to its architecture that the diff is directly applied on top of the existing filesystem, while in normal differential backups one might find out months or years later that one diff snapshot was damaged in transfer or is not accessible.
zfs scrubbing with S.M.A.R.T monitoring also helps a lot to ensure drive quality over time.
# Gotchas
ZFS:
- There is no undo etc, this is unix, so beware of wrong commands. - ZFS Scrubbing can be stopped (it does sometimes affect io speed), but ZFS resilvering cannot. This can lead to performance issues. - There must be enough RAM for the caching to work well and synchronous workloads do well with good write cache drives (ZIL) - Data usage patterns should fit well with the Append Log schema of ZFS. E.g databases such as LevelDB worked really well. Others are not slow, but need a good ZIL more then when the pattern fits.
SmartOS: Some minor gotchas with how vmadm deletes zfs filesystems or in general with SmartOS, e.g when having too many snapshots, but everything quite predictable.
Is nice to see it advancing in useful ways. Hopefully this leads to offline dedup without the runtime memory costs.
As I understand it, there will be no need to copy any data from the same dataset, and this includes all snapshots. Blocks written to the live dataset can just be references to the underlying blocks, and no additional space will need be used.
Imagine being able to continuously switch a file or a dataset back to a previous state extremely quickly without a heavy weight clone, or a rollback, etc.
Right now, httm simply diff copies the blocks for file recovery and roll-forward. For further details, see the man page entry for `--roll-forward`, and the link to the httm GitHub below:
--roll-forward="snap_name"
traditionally 'zfs rollback' is a destructive operation, whereas httm roll-forward is non-destructive. httm will copy only the blocks and file metadata that have changed since a specified snapshot, from that snapshot, to its live dataset. httm will also take two precautionary snapshots, one before and one after the copy. Should the roll forward fail for any reason, httm will roll back to the pre-execution state. Note: This is a ZFS only option which requires super user privileges.
[0]: https://github.com/kimono-koans/httmAll of this happens under the covers already if you have dedup turned on, but this allows utilities (gnu cp might be taught to opportunistically and transparently use the new clone zfs syscalls, because there is no downside and only upside) and applications to tell zfs that "these blocks are going to be the same as those" without zfs needing to hash all the new blocks and compare them.
Aditionally, for finer control, ranges of blocks can be cloned, not just entire files.
I can't tell from the github issue, can this manual dedup / block cloning be turned on if you're not already using dedup on a dataset? Last time I set up zfs, I was warned that dedup took gobs of memory, so I didn't turn it on.
>When --reflink[=always] is specified, perform a lightweight copy, where the data blocks are copied only when modified. If this is not possible the copy fails, or if --reflink=auto is specified, fall back to a standard copy. Use --reflink=never to ensure a standard copy is performed."
Also, as mentioned, on Linux, it's not wired up with any interface to be used at all right now.
For example, if you have a 1 GB file and you want to make a copy of it, you need to read the whole file (all at once or in parts) and then write the whole new file (all at once or in parts). This results in 1 GB of reads and 1 GB of writes. Obviously the slower (or more overloaded) your storage media is, the longer this takes.
With block cloning, you simply tell the OS "I want this file A to be a copy of this file B" and it creates a new "file" that references all the blocks in the old "file". Given that a "file" on a filesystem is just a list of blocks that make up the data in that file, you can create a new "file" which has pointers to the same blocks as the old "file". This is a simple system call (or a few system calls), and as such isn't much more intensive than simply renaming a file instead of copying it.
At my previous job we did builds for our software. This required building the BIOS, kernel, userspace, generating the UI, and so on. These builds required pulling down 10+ GB of git repositories (the git data itself, the checkout, the LFS binary files, external vendor SDKs), and then a large amount of build artifacts on top of that. We also needed to do this build for 80-100 different product models, for both release and debug versions. This meant 200+ copies of the source code alone (not to mention build artifacts and intermediate products), and because of disk space limitations this meant we had to dramatically reduce the number of concurrent builds we could run. The solution we came up with was something like:
1. Check out the source code
2. Create an overlayfs filesystem to mount into each build space
3. Do the build
4. Tear down the overlayfs filesystem
This was problematic if we weren't able to mount the filesystem, if we weren't able to unmount the filesystem (because of hanging file descriptors or processes), and so on. Lots of moving parts, lots of `sudo` commands in the scripts, and so on.
Copy-on-write would have solved this for us by accomplishing the same thing; we could simply do the following:
1. Check out the source code
2. Have each build process simply `cp -R --reflink=always source/ build_root/`; this would be instantaneous and use no new disk space.
3. Do the build
4. `rm -rf build_root`
Fewer moving parts, no root access required, generally simpler all around.
https://man7.org/linux/man-pages/man2/ioctl_fideduperange.2.... https://man7.org/linux/man-pages/man2/copy_file_range.2.html https://github.com/markfasheh/duperemove
Once this is ready, I am going to subdivide my user homedir much more than it already is. The biggest obstacle in the way of this has been that it would waste a bunch of space until the snapshots were done rolling over, which for me is a long time (I keep weekly snapshots of my homedir for a year).
In general, breaking up a filesystem into multiple ones in ZFS is mostly useful for making filesystem management more fine-grained, as a filesystem/dataset in ZFS is the unit of management for most properties and operations (snapshots, clones, compression and checksum algorithms, quotas, encryption, dedup, send/recv, ditto copies, etc) as well as their inheritance and space accounting.
In terms of filesystem management, there aren't many downsides to breaking up a filesystem (within reason), as most properties and the most common operations can be shared between all sub-filesystems if they are part of the same inherited tree (which doesn't necessarily have to correspond to the mountpoint tree!).
As far as I know, the major downsides by far were that 1) you couldn't quickly move a file from one dataset to another, i.e. `mv` would be forced to do a full copy of the file contents rather than just do a cheap rename, and 2) in terms of disk space, moving a file between filesystems would be equivalent to copying the file and deleting the original, which could be terrible if you use snapshots as it would lead to an additional space consumption of a full new file's worth of disk space.
In principle, both of these downsides should be fixed with this new block cloning feature and AFAIU the only tradeoffs would be some amount of increased overhead when freeing data (which should be zero overhead if you don't have many of these cloned blocks being shared anymore), and the low maturity of this code (i.e. higher chance of running into bugs) due to being so new.
But this is already a great advance, I love it :)
EDIT: Not to mention, this should make ZFS best-in-class for handling container filesystems. Deduplicating all the files in the individual layers means you can pretty much dispense with awkwardly trying to maintain inheritance trees.
So that syncs to my server, and my desktop/laptop. I drag files around on there when I want them deleted off my phone and archived somewhere. Me and my wife share a syncthing folder between us when we want to send files to each other.
All of this produces a fair number of duplicate files, particularly if you have backups turned on in case of deletes.
Offline dedupe basically makes all of that free though - duplicate files on the server or in backup dirs are no longer a problem.
Organizing that in shares means duplicating files...
Sadly I can't seem to find the presentation or recall the name of the project.
On the other hand, looking at for example RocksDB[2]:
File system operations are not atomic, and are susceptible to inconsistencies in the event of system failure. Even with journaling turned on, file systems do not guarantee consistency on unclean restart. POSIX file system does not support atomic batching of operations either. Hence, it is not possible to rely on metadata embedded in RocksDB datastore files to reconstruct the last consistent state of the RocksDB on restart. RocksDB has a built-in mechanism to overcome these limitations of POSIX file system [...]
ZFS does provide atomic operations internally[1], so if exposed it seems something like RocksDB could take advantage of that and forego all the complexity mentioned above.
How much that would help I don't know though, but seems potentially interesting at first glance.
Awesome!
https://blogs.oracle.com/solaris/post/reflink3c-what-is-it-w...
What's scary about it? You have to track references, but it doesn't seem that hard compared to everything else going on in ZFS et al.
> Is this common in filesystems? Or is ZFS striking out new ground here? At least BTRFS does approximately the same.
> What's scary about it?
It's scary because there's only one copy when you might have expected two. A single bad block could lose both "copies" at once.
3 Copies
2 Media
1 offsite.
If you follow that then you would have no fear of data loss. if you are putting 2 copies on the same filesystem you are already doing backups wrongs
You should also be able to restore your data in a calm, controlled, and correct manner. Test your backups to be sure they work, and to be sure that you're still familiar with the process. You don't want to be stuck reverse-engineering your backup solution in the middle of an already-panicky data loss scenario.
Remain calm, follow the prescribed steps, and wait for your data to come back.
Just that I'm trusting the OS to re-duplicate it at block level on file write. The idea that block by block you've got "okay, this block is shared by files XYZ, this next block is unique to file Z, then the next block is back to XYZ... oh we're editing that one? Then it's a new block that's now unique to file Z too".
I guess I'm not used to trusting filesystems to do anything but dumb write and read. I know they abstract away a crapload of amazing complexity in reality, I'm just used to thinking of them as dumb bags of bits.
likely zfs and btrfs have some similarities (copy on write) and some differences (integration with volume manager). if you're interested, i'm sure you'll find more.
This should end up being exposed through cp --reflink=always, so you could look up filesystem support for that.
1. Snapshot LVM volume
2. Mount the LVM volume and back the snapshot up somewhere else
3. Delete the snapshot as soon as possible
Otherwise you're going to lose a ton of performance.
[0] https://www.percona.com/blog/disaster-lvm-performance-in-sna...
But, it still has to duplicate metadata which depending on the amount of files may cause inconsistency in the snapshot.
The Arch wiki says:
"Tools dedicated to deduplicate a Btrfs formatted partition include duperemove, bees, bedup and btrfs-dedup. One may also want to merely deduplicate data on a file based level instead using e.g. rmlint, jdupes or dduper-git. For an overview of available features of those programs and additional information, have a look at the upstream Wiki entry.
Furthermore, Btrfs developers are working on inband (also known as synchronous or inline) deduplication, meaning deduplication done when writing new data to the filesystem. Currently, it is still an experiment which is developed out-of-tree. Users willing to test the new feature should read the appropriate kernel wiki page."
It's a similar effect only if you don't modify the files, I think.
If you "clone" a file with a hard link and you modify the contents of one copy, the other copy would also be equally modified.
As far as I understand this wouldn't happen with this type of block cloning: each copy of the file would be completely separate, except that they may (transparently) share data blocks on disk.