Or is there a better way to handle snapshots and data integrity (block level checksums)?
Without block level checksums, a single corrupted data block could corrupt half of your virtual machine images...
Or is there a better way to handle snapshots and data integrity (block level checksums)?
Without block level checksums, a single corrupted data block could corrupt half of your virtual machine images...
If ZFS gives you block level checksums, then you could use compression/dedup from VDO. Just activating i.e. compression on both layers would be waste of cycles.
ZFS dedup takes 320 bytes of RAM per unique block "record". So 1TB of RAM is enough to dedup only a bit over 6 TB worth of unique blocks, when using a block size that works well with virtual machines — 4 kB.
One can of course use larger ZFS record size than 4kB. But virtual machine dedup savings drop very sharply as a result if record size does not match virtual machine filesystem block size and alignment. This happens, because there are exponentially more different combinations how 4 kB blocks can be arranged inside a bigger record size.
It's painful to format all virtual machine images to use say 64/128 kB filesystem block size to be able to efficiently use larger ZFS record size dedup.
I understand VDO dedup requires significantly less memory and uses only 4 kB blocks for dedup. This is ideal for VM storage application.
LVM by itself doesn't do anything, it's a management layer on top of device-mapper. But yes, a block-level checksum exists since cryptsetup v2.0 (very recent), called dm-integrity: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
So you can set ZFS to dedup 4 kB blocks. It just requires a boatload of RAM.
This is one area where btrfs has a huge advantage, dedup can be done offline, rather than in real time.
My point is that the system will not stop working if the DDT is not in cache. It would still run, but it would not be fast.
But dedup is why I'm interested in VDO.
I've almost always left it at the default; on my home file server (slightly long-in-the-tooth 6 core xeon, enterprise-grade spinning rust), throughput is noticeably faster on compressible data.
(ZFS dedupe should only be considered for weird cases, like if you somehow have a ton of RAM but very limited storage. Frankly, at this point I think it is an attractive nuisance that leads beginners down a dangerous path and should be removed, or at least the commands to enable it should be given loud, scary confirmation messages.)
It's dm layers all the way down...
It’s documentation suggests that it can detect on journal replay whether the data was written where it was supposed to be, but it is not going to catch sites clobbered by a misdirected write every time because sone times the wrong sector is perfectly overwritten with no overlap with other sectors.
This is not a replacement for ZFS zvold, which do checksum each block.
I can't think of even one compression algorithm that implements checksums. That's why archive formats (like ZIP, 7z, etc.) need to implement checksums separately — inflate and deflate algorithms ZIP uses don't have any kind of built in checksums.
A block device layer would use the "raw" algorithm, not any frame/container format and use something like SHA256 for checksums.
Take a look at LZ4 source: https://github.com/lz4/lz4/tree/dev/lib
VDO LZ4 source:
https://github.com/dm-vdo/kvdo/blob/master/vdo/base/lz4.c
No LZ4F frame format or checksum in sight.
(by the way, sha256 is a slow cryptographic hash, not a checksumming one, like crc32)