I still have some boards with ~512mb RAM lying around (an UltraSPARC for example) that I'd love to re-purpose to a cheap NAS, just for the heck of doing it on a non-x86 platform....
I still have some boards with ~512mb RAM lying around (an UltraSPARC for example) that I'd love to re-purpose to a cheap NAS, just for the heck of doing it on a non-x86 platform....
ZFS dedup is block based, and actual block size varies depending on data feed rate for most workloads (zfs queues up async writes and merges them), so in practice once a file gets some non-zero block offset somewhere which happens all the time, even identical files don’t dedup.
My setup tries to get the absolute highest bandwidth and uses NVMe sticks in a stripe (I get my redundancy elsewhere), no compression, no dedup and yet can only hit ~ 3.5 GB/s reads (TrueNAS Core, EPYC 7443P, Samsung 980PRO, 256 GiB). I hope TrueNAS SCALE will perform better.
Though there is no clean way to disable either. Compression can be removed from files by rewriting them, but removing deduplication requires copying over all data to a fresh pool.
https://www.truenas.com/community/threads/zfs-dedup-disable-...
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 144M 14.9G - - 0% 0% 16.00x
ONLINE -
chicken:~# zfs send chicken_test/dedup_source@send | zfs recv -o dedup=off
chicken_test/nodedup_dest
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 2.29G 12.7G - - 0% 15% 16.00x
ONLINE -
chicken:~# zfs get dedup chicken_test/nodedup_dest
NAME PROPERTY VALUE SOURCE
chicken_test/nodedup_dest dedup off local
chicken:~# zfs destroy -r chicken_test/dedup_source
chicken:~# zpool list
NAME SIZE ALLOC FREE CKPOINT EXPANDSZ FRAG CAP DEDUP
HEALTH ALTROOT
chicken_test 15G 2.29G 12.7G - - 0% 15% 1.00x
ONLINE -One issue I had is that due to what I eventually tracked down as power issues, I had some corrupted data written to disk under my zfs pool (at the media write later), and I had dedup on.
So dedup, unfortunately, actually made it REALLY suck to fix, because I couldn’t even copy a new version of the file to the same pool! It kept nuking the duplication, and keep the old bad data and I then couldn’t read the copy. :s
It even did this after I deleted everything, because prune couldn’t remove the bad underlying entries because it was having a media failure.
So delete files, scrub, put new files on resulted in them having the exact same failure.
When I nuked the pool and recreated it, it was all fine though.
So yeah, be careful with dedup.
At the time, it was the difference between slow and impossible: I couldn't afford another 2x of disks.
These days, the pool could fit on a portable SSD that would fit in my pocket.
Careful, file-based dedup on top of ZFS might be more effective.
Small changes to single, large files see some advantage with block based deduplication. You see this in collections disk images for virtual machines.
You might see that in database applications, depending on log structure. I don't know, I don't have that experience.
For most of us, file-based deduplication might work out better, and is almost certainly easier to understand. You can come with a mental model of what you're working with, dealing with successive collections of files.
Even though files are just another abstraction over blocks, it's an abstraction that leaks less without the deduplication.
I haven't used a combination of encryption and deduplication. That was Really Hard for ZFS to implement, and I'm not sure how meaningful such a combination is in practice.
Hmmm, that 3.5GB/s sound low. From rough memory of doing initial storage benchmarking of our "new" Hetzner dedicated boxes a few months ago (AX51-NVMe, https://www.hetzner.com/dedicated-rootserver/ax51-nvme), they were giving about 10GB/s with mirrored NVMe drives.
Just logged into one of those boxes now, and it's running 2x 1TB Samsung PM9A1 drives (https://semiconductor.samsung.com/ssd/pc-ssd/pm9a1/), compression is on (lz4), and dedup is off.
(Didn't do any real tuning at the time, as these specs were already far in excess of what's needed for these servers.)
Would enabling lz4 compression be useful for your use case?
My understanding is that the OpenZFS project devs feels that the answer is an emphatic no, it should not exist, but they're committed to backwards compatibility so they won't drop it. (Take with a grain of salt; that's an old memory and I can't seem to find a source in 30s of searching.)
They discussed it, along with some options for a background-scanning dedup service (trying to find potential files to dedup), in the February leadership meeting[3].
[1]: https://openzfs.org/wiki/OpenZFS_Developer_Summit_2020_talks...
Why is dedup even present when the primary use case (backups) is better served in every way by snapshots?
In theory, it could also really help with virtualized disk workloads where there may be a lot of duplicated data from the base OS, but you can't use a snapshot (easily) because windows won't run from a zfs filesystem. You could maybe do snapshotting on zfs volumes, but that's not as flexible as a dedupe that worked as imagined.
Personally, I think online dedupe ends up being too expensive in memory and computation and ends up missing things because of divergent block sizes or offsets as another poster mentioned. ZFS doesn't support an offline dedupe, but I think btrfs does. That might be more interesting. It's still expensive to find duplicates, but it's possible, and it'd be neat to be able to rewrite the metadata to refer to a single copy and free some space, maybe.