A look at VDO, the new Linux compression layer
redhat.com
redhat.com
Could someone explain how VDO works to me? Based on the example, it looks like another DM backend, like LUKS and such. It exposes a virtual device backed by a real one, adding a layer of compression, am I reading it right? I can also see that we're specifying a "logical size" that is exposed to the user. How much space is really used and how is it allocated? Can I grow the logical size later?
Apart from this - what is the status of the patch, is it redhat-only or is it on its way to the upstream?
And deduplication, yes.
>Can I grow the logical size later?
Yes; the man page for the program discusses the growLogical subcommand briefly: https://github.com/dm-vdo/vdo/blob/6.1.1.24/vdo-manager/man/...
>I can also see that we're specifying a "logical size" that is exposed to the user. How much space is really used and how is it allocated?
There's a vdostats command mentioned in the article that exposes how much actual space is available and used.
>Apart from this - what is the status of the patch, is it redhat-only or is it on its way to the upstream?
It has not yet been submitted as a kernel patch.
1. Ran out of disk space,
2. Ran out of inodes,
3. Media reported ENOSPC and it gets propagated,
4. (something else I don't know about?)
So now, #3 gets more common. Is it actually something that FS devs think about? "What if I get this weird error message?" Or is it something that could trigger a some destructive edge case in the filesystem and lead to serious data loss? You mention dm-thin doing the same, so I guess that at least the popular filesystems should handle it well.
Anyway, when you think about it, how do you debug #3? You checked file size, you checked number of inodes, other than that the FS gives you no feedback. There's no central API to signalling this kind of problems and suggested solutions and I would say that this makes systems much less flexible. You could of course run vdostats (and whatever dm-thin uses to report its resources), but just imagine the amount of delicate code you would need to automatically solve this kind of issues. It's insane, it really looks like crappy engineering to me.
Filesystems (and VDO) both log extensively to the kernel log in an out of space situation, so inspection via dmesg or journalctl hopefully leads swiftly to identification of the problem. The 'dmeventd' daemon provides automatic monitoring of various devices (thin, snapshot, and raid) and emits warnings when the thin-provisioned devices it is aware of are low on space; there's a bug filed against VDO to emit similar warnings [1]. Careful monitoring is definitely important with thin provisioned devices, though.
I had a conversation with some of the RH kernel devs a few weeks ago and from my (hazy) memory, this comes from Red Hats acquisition of permabit so it was closed source, its now GPL'd but needs a fair amount of tidying up and bits rewriting to get it into a state that the kernel developers would accept it (they reimplemented some standard kernel features so they didn't have to link to GPL'd code). So it will be upstreamed but it will take time.
In the meantime its released for RHEL only, not even in Fedora. Iirc could be ported to other distros (its 100% GPLed) its just not been.
[1] https://rhelblog.redhat.com/2018/02/05/understanding-the-con...
At my employer, we're currently using btrfs in production for deduplication. It's relatively stable nowadays, but fragmentation and maintenance are massive issues for our use case.
_" The Btrfs file system has been in Technology Preview state since the initial release of Red Hat Enterprise Linux 6. Red Hat will not be moving Btrfs to a fully supported feature and it will be removed in a future major release of Red Hat Enterprise Linux. The Btrfs file system did receive numerous updates from the upstream in Red Hat Enterprise Linux 7.4 and will remain available in the Red Hat Enterprise Linux 7 series. However, this is the last planned update to this feature."_
Source: https://access.redhat.com/documentation/en-us/red_hat_enterp...
Works well but we'd rather not have to worry about these kind of things.
(even before it was deprecated, it was marked a "technology preview", rightfully so - if you want to use btrfs, you need to stay very close to mainline)
I have a long-term storage HDD that I thought I would never need data recovery on it until I fucked it up good by a wrong partitioning command. Normally those types of mistakes can be recovered by TestDisk but not that time. I realized my only hope was to use PhotoRec to recover what I can from the garbage that I overwrote on the disk. Thanks to PhotoRec I could recover most of them with messed up names. I was so thankful I didn't use any fancy compression algorithm.
Tools such as LVM, SSD backing storage, btrfs, compression and stuff are all nice when you understand their limitations. For now, I don't use all of my storage so I just create ext3, ext4, and exFAT partitions to store my data for the maximal chance of recovery, whether it is due to hardware or software.
Unless, of course, you have a ton of similar virtual machines or other data with high amount of 4 kB block size aligned redundancy.
Edit: Sorry not the disk image, the temporary overlay for zero blanking. But sometimes smaller disk images for the whole thing.
Recently I've also experimentally been holding some older games in RAM on a compressed image.
Or is there a better way to handle snapshots and data integrity (block level checksums)?
Without block level checksums, a single corrupted data block could corrupt half of your virtual machine images...
ZFS dedup takes 320 bytes of RAM per unique block "record". So 1TB of RAM is enough to dedup only a bit over 6 TB worth of unique blocks, when using a block size that works well with virtual machines — 4 kB.
One can of course use larger ZFS record size than 4kB. But virtual machine dedup savings drop very sharply as a result if record size does not match virtual machine filesystem block size and alignment. This happens, because there are exponentially more different combinations how 4 kB blocks can be arranged inside a bigger record size.
It's painful to format all virtual machine images to use say 64/128 kB filesystem block size to be able to efficiently use larger ZFS record size dedup.
I understand VDO dedup requires significantly less memory and uses only 4 kB blocks for dedup. This is ideal for VM storage application.
LVM by itself doesn't do anything, it's a management layer on top of device-mapper. But yes, a block-level checksum exists since cryptsetup v2.0 (very recent), called dm-integrity: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
So you can set ZFS to dedup 4 kB blocks. It just requires a boatload of RAM.
This is one area where btrfs has a huge advantage, dedup can be done offline, rather than in real time.
My point is that the system will not stop working if the DDT is not in cache. It would still run, but it would not be fast.
I can't think of even one compression algorithm that implements checksums. That's why archive formats (like ZIP, 7z, etc.) need to implement checksums separately — inflate and deflate algorithms ZIP uses don't have any kind of built in checksums.
A block device layer would use the "raw" algorithm, not any frame/container format and use something like SHA256 for checksums.
Take a look at LZ4 source: https://github.com/lz4/lz4/tree/dev/lib
VDO LZ4 source:
https://github.com/dm-vdo/kvdo/blob/master/vdo/base/lz4.c
No LZ4F frame format or checksum in sight.
(by the way, sha256 is a slow cryptographic hash, not a checksumming one, like crc32)
But dedup is why I'm interested in VDO.
I've almost always left it at the default; on my home file server (slightly long-in-the-tooth 6 core xeon, enterprise-grade spinning rust), throughput is noticeably faster on compressible data.
(ZFS dedupe should only be considered for weird cases, like if you somehow have a ton of RAM but very limited storage. Frankly, at this point I think it is an attractive nuisance that leads beginners down a dangerous path and should be removed, or at least the commands to enable it should be given loud, scary confirmation messages.)
If ZFS gives you block level checksums, then you could use compression/dedup from VDO. Just activating i.e. compression on both layers would be waste of cycles.
It's dm layers all the way down...
It’s documentation suggests that it can detect on journal replay whether the data was written where it was supposed to be, but it is not going to catch sites clobbered by a misdirected write every time because sone times the wrong sector is perfectly overwritten with no overlap with other sectors.
This is not a replacement for ZFS zvold, which do checksum each block.
The filesystem ontop always sees 3TB, unless you explicitly modify the VDO device. Of course, you have to monitor the VDO status tightly: if you happen to store data on the filesystem which is absolutely unique, has no zeros and is uncompressable, then dedup/compression can not do anything. The VDO device can only consume up to ~1TB of such data. Your monitoring should detect the low space on the VDO backend before, and you should then either stop writing or extend the backend device.
Should be tried out before relying on it. The current behaviour seems to be considered a bug, https://bugzilla.redhat.com/show_bug.cgi?id=1519377 has details.
I haven't bought a new hard drive in years, and I'm only using maybe 1.5 Tb of ~3 Tb, and most of that space is used up by RAW images from my DSLR. At the rate I'm going, it'll be 2 or 3 years before I need a new drive.