ZFS Deduplication
blogs.sun.com
blogs.sun.com
Shame ZFS has a slightly indeterminate future.
Solaris is doing poorly as a general purpose Unix OS to sell Sun support / hardware for, but as an embedded hardware/software appliance Sun kit is both better and cheaper than its counterparts, and they have some extremely talented engineers (though they've bled 27% of their staff since last year).
One word: 3Par.
(Longer version: 3Par was founded by ex-Sun types who wanted to make storage appliances based on Solaris. Sun responded by charging them outrageous licensing fees. They went with Linux and now their business is doing just fine...)
As an example, here's some people doing it right now: http://www.nexenta.com/
Edit: let me be more clear. When do you have situations where duplicate data is stored on the same disk and the best way to deal with it is through the filesystem?
If you're talking about a webapp with lots of users uploading the same photo or something, isn't that better handled before you hit the filesystem, so that you have dedupe over a number of independent disks/locations?
Database dumps as backups.
Email servers with lots of CCing.
Although it's probably more for 'Enterprises' than web apps.
But the great thing about this is that if operates at the block level - so if two people take a CAD drawing, change a small bit and save it to their home directories, most of the similar data can be stored once and only the changed blocks stored separately.
The real value (and the content of many EMC sales pitches) is with backing up and archiving data. EMC often argues that most "corporate" data seems not to change very much over time.
I don't see how version control avoid file duplication. When you work on multiple branches at the same time. You necessarily have to retrieve each branch in different directories.
E.g., if you store dated snapshots of content, it's much easier on you if you can just store the whole directory, rather than manually maintaining hard links back to previous versions.
In my toy backup app I'm doing pretty much what they're doing - I assume that each block with 1MB of data will have a unique hash. I haven't tested it on serious amounts of data though, and so I'm very curious to know how likely this scheme is to survive an encounter with the real world.
If you hash 4KiB-long blocks, then every possible block will share the hash value with (on average) 128 different 4KiB blocks. And on a standard 200GiB disk you can fit (more or less) 52,400,000 blocks.
This explanation is a bit less reassuring. Now consider the fact that your data is never random and you hit the same patterns all the time (loads of zeros / ascii letters / x86 code)
You can also be reassured that there's lots of research going on about how likely these collisions are and how to find them. People are actively trying to break these hash algorithms, so it's not just in theory.
My initial remedy is to add another hash method and name the data based on the results of both; that way problematic data would need to trigger a collision in two different algorithms at the same time, which "should" be next to impossible. Currently my file fragments are named "{SHA256}.dat", in v2 I could instead name them "{SHA256}{SOME_OTHER_HASH}.dat".
Web archiving is another application that can benefit from this. You can crawl the same website 10 times a day and just store all of the files as it is.
Incidentally, ZFS also has an option to compress the data. You have a choice of a fast but not so wonderful algorithm, or gzip level 1 to 9. Since de dupe is at the block level, I believe you can combine the two.
While it might be true for some compression formats/programs, it is not the case for .tar.*, as the directory is archived to the single file first (tar), and then compressed. So if you have similarities within 2 different files, that will be exploited.
I think to make something work for "large blocks" that already does a good job for "small chunks" is just finding the right values for parameters of compression algorithm used.
tar.x will compress chunks of data that is smaller than zfs block size. ZFS dedup + ZFS compression will compress them as well, so what's the problem?
Imagine you have 10 VMware virtual machines, and all the VM files are stored on a ZFS server with dedupe support.
Let's say that the base install image is 2GB in size of OS, GUI, Java libraries, etc.
Non-dedupe scenario means 20GB of storage is used.
Dedupe scenario means maybe, 2.5GB of storage is used due to slight differences in the way the blocks are arranged, etc.
Now you add an 11th VM, same OS, same base setup. Storage goes up by maybe 100MB even though the VM has the same 2GB of files.
Do you see how zipping a file would not show you that effect?
On HN, it is downvoted and reply with technically wrong claim gets 4 points.
And, BTW, I am in no means implying this is bad. Quite the contrary - I would love to have an inexpensive SPARC-based desktop. If it existed.
CRC32 is faster than ZFS's default of Fletcher2 and has less frequent collisions.
While it might seem flippant to compare VAX to x86, it was found (on VAX) that a programmer could potentially get better aggregate performance by avoiding that instruction; that having an instruction doesn't automatically mean that the code is faster.