Initial Hammer2 filesystem implementation
lists.dragonflybsd.org
lists.dragonflybsd.org
- This is DragonflyBSD's next-gen filesystem.
- copy-on-write, implying snapshots and such, like ZFS, but snapshots are writable.
- compression and de-duplication, like ZFS
- a clustering system
- extra care to reduce RAM needs, in contrast to ZFS
- extra care to allow pre-allocation of files by writing zeros, something that will make SQL databases easier to run performantly on HAMMER2 than on ZFS
And much more. The design doc is an interesting read, take a look:
https://gitweb.dragonflybsd.org/dragonfly.git/blob_plain/HEA...
Isn't that an oxymoron? I thought the entire point of a snapshot was that it's an immutable record of the filesystem at a moment in time.
extra care to reduce RAM needs, in contrast to ZFS
ZFS utilizes a lot of RAM for caching. That's not the same thing as needing a lot of RAM. I've seen this same complaint about modern operating systems. People will buy a lot of RAM and then get annoyed when they see the OS making use of it. As long as the memory is available to be allocated to applications, why should we care whether or not the operating system makes use of it?
I believe the complaint about ZFS and RAM is that the caching cannot be considered optional. Performance is substantially worse than other filesystems with more modestly sized caches, so the memory isn't really available to applications.
https://docs.oracle.com/cd/E53394_01/html/E54818/chapterzfs-...
Is that in comparison to other filesystems that provide the same level of data integrity as ZFS? Otherwise you're comparing apples and oranges.
1) data integrity
2) low memory use
3) performance
For comparison:
ZFS has an online-only dedup, so it can save space as data is written, but it can't combine identical pieces of already-written data. It also scatters the segments of files to the winds, needing lots of ram and fast disks or it slows to a crawl.
BTRFS has mostly-offline dedup, so it doesn't save space as data is written, but you can combine identical pieces of data later. You can also make CoW copies of files instantly. It has minimal performance impact.
There are pros and cons to each approach, and as usual with storage, a lot depends on implementation and workload.
Thanks.
> There are pros and cons to each approach, and as usual with storage, a lot depends on implementation and workload.
I would probably argue that the biggest factors in different performance between ZFS and BTRFS deduplication are nearly independent of when the deduplication happens, and boil down to implementation decisions specific to those two filesystems.
For example some filesystems can be grown online, but can shrink only offline.
But this one seems weird to me:
> copy-on-write, implying snapshots and such, like ZFS, but snapshots are writable.
snapshot by definition should be read-only, also ZFS allows to create a new filesystem from snapshot (through zfs clone) which is also O(1) operation.
One reason why you would want a writable snapshot is during backup validation. An application may need to do crash recovery before being able to read data, and having a writable filesystem makes things a lot easier.
I know that ZFS and LVM both send writes to a separate area, so the original snapshot is never modified (don't know about Btrfs implementation details).
Especially if the default is to have a read only snapshot, I do not rule that out.
I always though it care about performance of SQL databases you bypass the filesystem / kernel and go directly to the LUN with something like ASM.
Does it support RAID and subvolumes like Btrfs? How is the stability?
> Are there any preliminary benchmarks available?
He spent years just to write the doc and spec. And the post just said the preliminary code will be in September so that's probably a no.
Even with the september code being posted, preliminary isn't going to mean much because it most likely have very little features compare to ZFS and Btrfs. You probably have to wait longer for a fair comparison...
IIRC it's a one man team and the dude is a unicorn.
IIRC, years ago Dillon sold some proprietary synchronous multi-master replication product that provided ACID guarantees while also being relatively performant. (Because synchronous multi-master had historically been quite slow.) I always thought HAMMER2's replication model was going to be an evolution of that tech.
I really hope Dfly gets more adoption and broader use.
If you accept the mathematical claim that a secure hash like SHA256 has
only a 2\^-256 probability of producing the same output given two different
inputs, then it is reasonable to assume that when two blocks have the
same checksum, they are in fact the same block. You can trust the hash.
An enormous amount of the world's commerce operates on this assumption,
including your daily credit card transactions. However, if this makes
you uneasy, that's OK: ZFS provies a 'verify' option that performs
a full comparison of every incoming block with any alleged duplicate to
ensure that they really are the same, and ZFS resolves the conflict if not.
To enable this variant of dedup, just specify 'verify' instead of 'on':
[1] https://blogs.oracle.com/bonwick/zfs-deduplication-v2E knows of some critical security update that needs to be installed in sensitive locations.
E also knows of some attack on the hashing algorithm that is in use by the filesystem to craft a small block containing mostly garbage but some key bits that they would like to control. (Yes this is the hypothetical, but prior algorithms /have/ fallen).
E thus arranges to have this 'duplicate' block stored before routine and predictable maintenance patterns.
A installs the updates and the 'duplicate' file is now E's datastream, but A's intended credentials.
E has caused system corruption, and potentially privilege escalation.
512 bit hashing is basically placebo.
Thus even with FreeBSD a port would not be straight forward because of VFS API differences, required scheduler support and buffer cache implementation differences[1]. Linux I would assume is even more work.
I started an improved port of HAMMER to Linux (I don't like the current FUSE one and started building an in-kernel version from scratch) but, stopped working on it before I got it working or published it. I am planning to start working on it again with someone else. I wouldn't be surprised if someone else beats us to it though.
What utilities you need and, what you do in userspace, depends on what the goal of the port is somewhat I think. Also, what goes in user-space depends on what you are porting to (FUSE or what OS). My goal was (and is again now) to have first class support on Linux, i.e. you should eventually be able to do everything you can on Dragonfly with HAMMER on Linux, with as good performance as possible. Initially, performance takes a back seat to feature support (unless it were to be so bad that it made features unusable) but, good performance is an important end goal of mine. I think for my goals, an in-kernel driver and user space utilities based on the Dragonfly ones but, with some substantial differences in code, make sense.
There is a FUSE port of HAMMER that lets you read, without history I think. I don't think this is maintained anymore but, it had very different goals to mine and, a port with different goals to my own, might want to pickup this/start a new FUSE port. In particular, if you don't care about supporting every feature, you want to get things that you do plan to support working more quickly, you don't care about performance in the long-run and/or, you want to target other platforms than Linux as well/instead, FUSE might make more sense. I think there are reasons to work on a FUSE port as well as a specific port to Linux (and other OSs), they have slightly different goals and trade-offs.
To be honest, I'm not sure exactly what you want to know here so, if I haven't addressed it, please do ask (but, it is possibly the case that what I know is too out of date to be interesting).
If you're interested in working on this together (there are a few people I know already who have expressed interest in working on this, all based in London so far), hearing more about progress in the near future or, if you are planning your own Linux port or port to another OS, I'd be interested to hear. Let me know if you'd like me to provide contact details for non-public messages (there's no email on my profile right now). I am not planning to make anything I do public until at least we have read support, with no history, working and at least somewhat tested but, I'm not set in stone about that particularly. I don't see any reason to use our port over the existing FUSE one until after then either.
What would you say to me or show to me to sell me on the idea of using a BSD?