Raid-Z Expansion Feature for ZFS Goes Live
freebsdfoundation.org
freebsdfoundation.org
Always reminds me of when NetApp used to do their arrays in RAID-4 because it made expansion super-fast, just add a new zeroed disk and only had to update the new disk blocks + parity drive on writes. Used to blow our Netware admin's mind as almost nobody else ever used RAID-4 -- I had it as an interview question along with "what is virtual memory" because you'd get interesting answers :)
After that, can we have defragmentation?
Let's make the best of both worlds.
I can see a conceptual sort of super-scrub that balances a zpool and addresses all that that's not on anyone's radar AFAIK.
Workload is roughly
Every night write a lot of files
6/7 nights delete a lot of files (i.e. keep weekly snapshots).
It took about 3-5 years of doing this to get to the point where write-performance dropped off the clif and enabling metaslab_debug brought it back.
a) random seek/many-small-files performance has been really had since day 1, I initially suspected old/low-end hardware (i3, 1600MHz RAM) but given that I can do just south of 200MB/s (two-way mirror) I'm kinda staring at ZFS expectantly here
b) I've admittedly managed to net myself a fair few pathological way-too-many-files situations from projects and whatnot that I really do need to get to cleaning up
Fairly early on I noticed apt performance degraded pretty badly, and long before (b) became a substantial concern it got to the point where installing just about anything would take about 60 seconds to do the "Reading database ..." step.
I've been idly curious about tweaking different settings to try and improve performance, but it's mostly been an idle curiosity because I don't have a straightforward way to back out of "oh great now what" edge cases.
This has probably been going on for just around a year or two, and with absolutely no context I'd be confident saying write volume isn't a shadow of what you're doing :) so perhaps that particular tunable is... maybe not relevant? Or maybe it is. I'm curious.
“i3, 1600MHz RAM” sounds like a laptop. Are you doing anything funky like using USB HDD enclosures?
Also try comparing the number, size, and latency of IO operations submitted to ZFS vs the same stats for IO submitted to the disks with https://github.com/iovisor/bcc
Once you figure out what layer (application? VFS/cache? file system? IO elevator? HBA? disk firmware?) the performance drop is happening on, it should be trivial to fix.
It's not a laptop, it's a low-end motherboard currently serving as my primary workhorse :) (until I find the money to fix the issues preventing me from working... any day now... :'D). *Checks* It's an ASUS P8H61-M. And no, the disks are directly attached.
TIL ZFS can submit different IO sizes than what reach the disks. I've just been dumbly staring at iotop and thinking that was the last word on the situation. Now to figure out how to get that info from ZFS (and figure out which bcc script to use). Thanks.
Thanks for the layer consideration. The application layer (an ncdu scan I'm currently doing is has been reading 60 files/second for days) and VFS/cache layer (if I do two apt operations in relatively quick succession (seconds apart) with nothing else doing I/O, the second one completes the read step instantly) seem to be the effect/symptom, with file system (all ZFS, but obviously badly tuned) and IO elevator (oooooh that's what that is TIL, I might play with this! :D) seemingly the most interesting, and HBA (onboard SATA3 port *hides*) and disk firmware (I've never upgraded a BIOS in case I irreparably break something lol) beyond the horizon somewhat.
Thanks for the info!
A sibling mentioned making sure ashift is 12. I'll second that. In addition, make sure your ZFS partition is aligned; if you gave ZFS the whole disk it probably is. If you did not (e.g. because you needed an EFI boot partition) it might not be.
Lastly, for any given workload, ZFS seems to have roughly logistic performance curve with respect to the amount of RAM it has to work with. The ARC does a pretty good job of keeping important data in RAM to minimize seeks when there is "enough" RAM, but it does a progressively worse job as it gets RAM constrained. On a development machine where I'm dealing with multiple SVN and git checkouts on spinning metal the performance difference between 8GB and 12GB of RAM for ZFS is night and day. Good SSDs make this a lot less important because the penalty for a small number of read-misses is approximately zero compared to rotating drives.
ashift is definitely 12 (as I noted in my sibling reply)... but TIL about alignment (thanks). Parted says my ZFS partition starts at 8590983168 bytes (after an 8GB swap partition), which divides down by 4K cleanly. Is that what you mean?
Hmm, the RAM usage on this machine is generally low-ish, but with 8GB I suspect the smallest perturbations can make a big difference (even though I use it headlessly). I'll definitely keep more RAM in mind going forward, and yeah, SSD/NVME storage makes these kinds of considerations moot in high-performance contexts.
Thanks for the info!
With XFS, I can xfs_fsr to undo it. With Btrfs, I can fi defrag to undo it.
With ZFS, reformat the disk. I'm not comfortable with that.
https://www.usenix.org/system/files/login/articles/login_sum...
As a home server administrator, I've wanted this feature for so long. Before this, in order to expand an existing array I'd have to fail every single drive and replace them with new higher capacity ones.
The only question I have is whether it supports expansion with drives of different capacities.
I’m also very interested in this - my main reason for sticking with btrfs is that I can use a variety of odd-sized drives, and expand it by adding a new oddly-sized drive...
RAID5 has the benefit of not making the parity drive a bottleneck that will fail sooner.
The code is here[0]. It still needs more testing and cleanup, and will then eventually be merged. After that it'll take some time to make to all the distributions ( freenas, freebsd, etc. )
> But the interesting thing about this project is how did it come to be. So a long-requested feature - how did it get funded? So actually, it’s funded by the FreeBSD Foundation.
Now I found why OP posted by FreeBSD Foundation.
One downside that I see of this approach, if I understand it correctly, is that the data already present on disk will not take advantage of the extra disk per slice. For example, if I have a raidz of 4 disks (so 25% of space "wasted"), and add another disk, new data will be distributed on 5 disks (so 20% of space "wasted") but the old data will keep using stripes of 4 blocks, they will just be reshuffled between the disks. Do I understand it correctly ?
Rewriting data in place can be tricky, if you have old enough snapshots the newly added space may not be enough. I hope they will find a good enough solution.
The hierarchy is disk < vdev < zpool.
Disk is physical.
vdev is logical. Purpose: Disk grouping and redundancy. Composition: One or more disks.
zpool is logical. Purpose: Higher-level management of one or more vdevs. Composition: It acts like a JBOD.
---
zpools can be thought of as "stripes of vdevs". This, in the narrow sense that the failure of any vdev in a zpool is a permanent loss of the entire zpool. All your redundancy in the ZFS ecosystem is via mirrored or RAID'ed vdevs.
---
The setup I have heard of that balances performance, redundancy and space is to do what you say: Have a zpool of multiple mirror-type vdevs.
You can also stripe at the vdev level, which I would assume has higher performance than having multiple single-disk vdevs in a pool - I'm unaware of the differences at a low level.
Naively I would expect that to become 25 data writes per disk, and then the fsync would go to all disks.
That might be about 30 writes total, so I'd expect it to be at least 3x as fast as a single disk.
Where does that expectation break down?
The 500 disk writes are parallelized over 5 disks so they only take the time taken for 100 writes (500 / 5).
So the IOPs is the same as a single disk.
HOWEVER, bandwidth is 4x. The above assumes that seek time dominates. If your writes are multiple slices in length, they will be written 4x faster because the amount of data per disk is divided across the disks. If you're reading and writing large contiguous files, then you do get a big I/O boost from raidz.
And if I wrote 100 slices to a single drive, is that 100 writes or 400 writes?
If it's 100, then it sounds like the slices are sized wrong: they should get bigger when I add more disks.
If it's 400 writes, then 100 writes per disk should be much faster.
For writes, the system is copy-on-write, so for each write, there is a full read of the containing record, often across all disks, so the updated record can be written out in the new location on the disks without disrupting the old record, which may still be referenced by snapshots, etc.
Adding additional disks gets you more throughput (you can split a single large write across more disks, but not more OPS.)
There are ways to work around elements of this and get much better performance, but not wanting more complexity makes enough sense.
Definitely a flaw in the ZFS data model though, rather than something inherent to the use of multiple disks.
It's a reasonable choice, but only because it makes certain kinds of bugs harder, not because it's safer when the code is correct.
> Any error that may happen during would destroy that data - while without moving it will likely be recoverable even with a dead harddrive.
That's not true. You make the new copy, then update every reference to the new copy, and only then remove the old one. If there's an error halfway through then there's two copies of the data.
That is just maths.
If you want to use ZFS and have more IOPS you will have to have more VDEVs (either more RAIDZs or several mirrors). Your storage efficiency will be slightly reduced (RAIDZ) or tank to 50% (mirrors) with less redundancy.
But to call it a flaw in the data model goes a long way to show that you do not appreciate (or understand) the tradeoffs in the design.
You could have both! That's not the tradeoff here. The problem is that if you wrote 4 independent pieces of data at the same time, sharing parity, then if you deleted some of them you wouldn't be able to recover any disk space.
> That is just maths.
I don't think so. What's your calculation here?
The math says you need to do N writes at a time. It doesn't say you need to turn 1 write into N writes.
If my block size is 128KB, then splitting that into 4+1 32KB pieces will mean I have the same IOPS as a single disk.
If my block size is 128KB, then doing 4 writes at once, 4+1 128KB pieces, means I could have much more IOPS than a single disk.
And nothing about that causes a write hole. Handle the metadata the same way.
ZFS can't do that, but a filesystem could safely do it.
> But to call it a flaw in the data model goes a long way to show that you do not appreciate (or understand) the tradeoffs in the design.
The flaw I'm talking about is that Block Pointer Rewrite™ never got added. Which prevents a lot of use cases. It has nothing to do with preserving consistency (except that more code means more bugs).
It's reasonable to expect that issuing a batch of 100 writes at the application layer followed by a fsync would not always require doing 100 writes to each underlying block device. The OS/FS should be able to combine writes when the IO pattern allows for it, and should be doing some buffering prior to the fsync in hopes of assembling full-stripe writes out of smaller application-layer writes.
It sure shouldn't be. This isn't dumb RAID.
I'm surprised that this would be a little-known caveat; it's been a problem with all RAID-parity systems 3,4,5 and Z) since forever.
It's a constraint on iops in particular, not bandwidth, because writing a record to disk involves writing something to every disk in the pool. So if you have a disk that's taking a long time to write, the other disks need to pause and let it catch up.
The old way ( as you referenced ) was to replace each disk one by one with a larger one.
If I'm understanding this right, and please correct me, this feature will allow you to add a 5th disk to a 4 disk raidz
And if I'm right about that, then this feature wouldn't really make sense for RAID10 anyway
Adding drives to a mirror has worked in zfs since prehistoric times. “zpool attach test_pool sda sdc” will mirror sda to sdc. If sda was already mirrored with sdb, you now have a triple-mirror with sda, sdb, and sdc.
You probably already know this, but btrfs offers a different option. Its "RAID1" mode would not result in the three drives storing your data in triplicate but rather still stores data in duplicate but with a usable capacity that's 1.5x that of a single driver's capacity. To get the ZFS behavior you describe, btrfs offers a "RAID1c3" mode, but that's a much less interesting feature.
Asymmetric RAID is probably the only config that I'd like to see in ZFS - being able to have (say) an 8TB drive and two 4TB drives, and have 8TB of mirrored space is quite nice when you're just buying commodity kit and don't want to have to retire/match drives all the time when growing mirrors.
Your example doesn't expand the storage of the vdev, which is what this entire discussion is about, it merely adds a mirror.
When you want to add a single drive to expand your available storage capacity, without reducing the degree of redundancy you're already using. That use case should have been obvious given how much discussion it's already received in this thread, so I think you have some kind of blind spot about anything small-scale. This is why it's good to actually understand the competition, even after deciding which tradeoffs are right for you.
This is still answering the wrong question, unless by "add mirrors" you mean "add another drive to a mirrored vdev and get more usable capacity", which I don't think is how ZFS works and would be strange to summarize as "add mirrors".
I'm trying to discuss an idea you don't seem to have any terminology for. Can you please try to meet me halfway and at least respond in a manner that makes it clear you understand there's a real distinction to be made?
Forcing ZFS to store multiple data blocks also can be done.
This is great for ZFS on flash drives.
https://docs.oracle.com/cd/E19253-01/819-5461/gevpg/index.ht...
AFAIK it's being developed against FBSD, but is not any more integrated there than it is in the main OpenZFS project.
If you want "feature chasing", FreeBSD ZFS TRIM is the ur example. I've read that code end to end... and I'll leave it at this.
Of course, it gives "better publicity" for RAID-Z, but it is rather an marketing trick not engineering.
In other words, what they are saying is that the traditional way to expand an array is essentially to rewrite the whole array from scratch, so if the old array has three stripes, with blocks [1,2,3,p1] [4,5,6,p2] and [7,8,9,p3] (with p1 and p2 being the parity blocks), the new array will have stripes [1,2,3,4,p1'], [5,6,7,8,p2'] and [9,x,x,x,p3'], i.e. not only has to move the blocks around, but also recompute essentially all the parity blocks.
IF I understand the ZFS approach correctly, the existing blocks are not restructured but only reshuffled, so the new layout will be logically still [1,2,3,p1] [4,5,6,p2] and [7,8,9,p3] but distributed on five disks so [1,2,3,p1,4] [5,6,p2,7,8], [9,p3,x,x,x]
It seems that this means less work while expanding, but some space lost unless one manually copies old data in a new place.
IF I got it right, I am not sure who is the intended audience for this feature: enterprise users will probably not use it, and power users would probably benefit from getting all the space they could get from the extra disk