This allows the dedup table (which stores the hashes of existing data to find identical blocks) to be limited in size to whatever you want, and especially prevent it from eating all the RAM. The tradeoff is that the smaller it is, the deduplication hit rate drops.
Here's a talk by Matt Ahrens about it, which also includes some other tricks to improve dedup performance:
As a result ZFS with deduplication on uses a ton of ram and has slow writes. This is so big an impact that most folks just turn it off, despite the tempting lure of "free" disk space.
The current implementation of dedup in zfs was the biggest disappointment I had in setting up my current storage box. Even with plenty of horsepower, I quickly turned it off and accepted the net reduction in capacity.
[1] I put the wrong command here originally.
Or actually detects duplicate blocks? Because that would be still be something, knowing how much deduplication would benefit you, if implemented.
Any filesystem with data checksumming should only need to build an index of the checksums to find duplicates, which sounds far cheaper. Does btrfs do that?
The utilities can be written in any language, there's several popular ones called dupremove, bedup, btrfs-dedupe, bees, and dduper.
Bedup specifically mentions "It integrates deeply with btrfs so that scans are incremental and low-impact."
For my use I'd likely ignore all files less than a few weeks old for deduplication purposes. Any virtual machine images I'd use block level deduplication and everything else I'd use file level deduplication.
Doesn't seem much different overhead wise than the normal scrub that linux software RAID, ZFS, and btrfs does anyways. I'm all for not storing old files twice, but would like to avoid the overhead of checking for deduplication on every write.
If they did, the dedup process for a multi tb drive would just be a few seconds, and the overhead would be so small it could be run on the whole drive every minute all day long.
Yes you can deduplicate 1TB quickly if you already have the data and metadata data checksums.
ZFS (with dedupe on) checks the checksum of every written block against every existing block in the pool, thus the slow writes and being very ram intensive.
Btrfs instead calculates data and metadata checksums for integrity then allows using them afterwards for deduplication. You could of course check for block level checksums for 100% of your storage. Or you could decide to dedupe only old files, only VMs, or only non-VMs. Or even a hybrid of block level for VMs, but file level for non-VMs.
Lots of software is not as CoW-friendly as it could be, offline dedup fixes that with relatively low overhead.
Running a dedupe based on file rather than block would do the same thing. You don't gain much if anything in your example using btrfs.
What is being argued here?
* fdupes will let you make hard links on any filesystem, but hard links do the wrong thing when you edit them or need different permissions
* fdupes will not let you reconcile duplicates in snapshots on ZFS
* ZFS deduplication is block-based
Hard links mostly work as you tell them to. You may need to set things up properly, but it's essentially the same thing being done.
>fdupes will not let you reconcile duplicates in snapshots on ZFS
Why not? The files are kept in a hidden folder in the base directory.
>ZFS deduplication is block-based
It is also live. You don't gain Mich with offline block deduplication over file, though it is better. Live deduplication is a feature even if unnecessary for most.
For files that are used by anything, they might be written to. I can't have a write to one file appear in a completely different file just because they had the same contents at one point. And what if my duplicate files have different owners or permissions?
> Why not? The files are kept in a hidden folder in the base directory.
Those are read-only views, not real folders.
> You don't gain [much] with offline block deduplication over file, though it is better.
In general, perhaps not. But for ZFS specifically you gain a lot by not using its deduplication system. It needs tons of memory and explodes files into fragments.
The act of saving should destroy the link, though I know that isn't always the case. Still, this is a benefit of a CoW filesystem, not offline dedupe. Permissions may be a problem, I'm not great with them.
>Those are read-only views, not real folders.
Then mount them writable,
>But for ZFS specifically you gain a lot by not using its deduplication system. It needs tons of memory and explodes files into fragments
Live deduplication is valuable in some instances, otherwise there are still offline alternatives available.
It's one method of gaining CoW. And ZFS does not support CoW between files. But whatever feature you lump it under, it's an advantage Btrfs has.
> Then mount them writable,
You can't.[1] And even if you could, you can't make a hard link between subvolumes.
[1] You can 'clone' a snapshot to make a writable subvolume, but the original read-only snapshot can never be deleted while the clone exists, so this doesn't help you fix duplicates.
That's bullshit. hard links are the same exact inode, with all that entails. reflinking is at the data block level. Totally different.
Reflink has its own fs metadata including inode, with (initially) shared extents. Those shared extents can have their blocks individually and independently modified, per file. The point at which there are no more shared blocks, they're not reflinks.
It's not an intrinsic property of file-dedupe vs block-dedupe. It's just how it's conventionally done.
Inodes that happen to share blocks are not the same file. ie totally different.
Of course this only works if the data is read-only.
It won't reclaim any space on ZFS if you're using snapshots.
I think that block-level deduplication has more predictable performance though: for example if you write some blocks to a deduplicated file, with block-level deduplication you just write the new blocks, with file-level deduplication you have to rewrite the whole file.
I do use zfs in preference to ufs on my pfSense boxes running non RAID on mSATA type storage but that is not a scientific choice but based on some quite iffy hearsay.
I have my opinion and you have yours. When it comes to filesystems then I want simple (and good backups.) Fast is nice but reliable is better. If I want faster I buy a bigger RAID controller or whatever.
Their backup options are drastically more complex because there's no way to take an atomic snapshot.
The complexity of data-integrity and backups are pushed a layer up, whereas more complex filesystems manage data-integrity and atomic snapshotting themselves, and can do so more easily due to existing at the correct layer.