OpenZFS – dRAID, Finally
klarasystems.com
klarasystems.com
[0] https://docplayer.net/117362530-Draid-declustered-raid-for-z...
I believe Windows Storage Spaces has long used a similar technique[0] too, though it has been a few years since I was looking into it. I wonder if ZFS DRAID will eventually allow full utilization of disks of mixed sizes too, that was always a very attractive feature of Storage Spaces to me.
[0] https://techcommunity.microsoft.com/t5/storage-at-microsoft/...
The numbers are not explained, nor are the rows, or columns. And the numbers repeat.
I saw the diagrams first outside of that pdf, though. Still, the way it goes "4, 8, 2, 16, 1, and so on..." is a boatload of "wat?".
So the RAIDZ is set up with fixed groups of drives, but the dRAID is not, even though the parity ratio is the same.
Your comment explains it better than the illustration.
> The dRAID offers a solution for large arrays, vdevs with fewer than 20 spindles will have limited benefits from the new option. The performance and resilver result will be similar to RAIDZ for small numbers of spindles.
If you're doing a 6-8 disk array, this probably isn't relevant to your interests.
If you're doing a 20+ disk array, well firstly, damn, you've been lucky if you've avoided any serious issues. Secondly, you really want to look into this ASAP.
Then you might want to consider ceph.
If you're comfortable with the node being a single point of failure, you can do a lot with a single storage node connected to a JBOD and serving data over NFS. And even if the node is the unit of failure, you can attach a SAS JBOD to multiple servers if you'd like.
In my position, in addition to the added cost, I just don't have the physical rack space for the extra servers required for ceph.
Sure, it depends but the 140 drives (mentioned by secabeen) alone will take up a huge amount of space. Maybe I just wanted to point out that there might be a point when you have maxed out the storage capability of a single node and might want to consider a distributed storage setup.
Right now, I use ZFS with multiple raidz2 or raidz3 vdevs with a number of hot spares waiting to take over. Resilvering can take a day for a 4TB disk. I’m terrified how long a 12TB disk will take (hence the move to raidz3). Disk loss is a normal fact of life, and it’s something we prepare for. But this would help me sleep better at night.
I’m seriously excited about this and I’m definitely the target audience for this.
Remember - if you are in the middle of a potentially failing array, every single bit of parity data is pure gold and if you fail out a drive, that drive worth of parity is gone forever.
So if you fail out a drive but it was actually weird bus errors getting spammed by some other drive that is actually failing, now you're in trouble ...
While on the topic, in filesystem emergencies, don't forget the ability to mount, and even run, read-only. Your filesystem might be puking on mount and backing you into a corner, but perhaps if it were mounted read-only ...
Which is to say, beat the shit out of the disks before deploying them[1] to weed out early failures, then aggressively fail them out at the first sign of trouble[2].
We don't use hot spares - we just very aggressively pull drives that give us any meaningful smart errors ...
[1] perhaps with the badblocks tool ...
[2] smartctl
- IBM's GPFS
- Panasas' PanFS
- Xyratex's -> Seagate's -> Cray's -> HPE's GridRAID in Clusterstor/Sonexion prodcuts
Though ZFS and dRAID will likely show up in that last one in the future.
As lustre is mostly a network raid0 (its got better now) it makes sense as you really shouldn't be using it for unreconstructable data.(gross simplification alert)
This is because you need to do a coordinated read from each disk in the raid stripe by stripe. a rebuild is effectively a single threaded synchronous operation
We don't really need to have such rigid mapping of data to raid groups anymore. which is why most large storage systems partition your data up into small chunks (a few megs) apply some forward error correction to it, and smear it all over the place.
This means that you can do multi-threaded rebuilds. This means that as your array gets bigger, not only does the rebuild time get smaller, but its much more resilient to data loss.
A good example of this is GPFS's native raid: https://www.usenix.org/legacy/events/lisa11/tech/slides/deen... (https://www.youtube.com/watch?v=2g5rx4gP6yU)
A word of warning though. Don't do this for small arrays (<50 disks) its not worth the hassle, yet.
After diving into consistent hashing I was surprised that all blocks are striped across all disks instead of redundancy+1,2,3 disks, in pretty much any drive array solution I could find architecture docs for. Seems like bigger blocks on fewer disks would have better utilization numbers. You could be doing 2 unrelated reads at the same time from different spindles and tracks.
With a dead drive, you could do nothing and wait for a new drive, or start making backup copies across the rest of the array as if the array were now n-1 disks, or some of both, like this solution seems to do.
Didn't sequential resilver already solve the read half of the problem? Because that would already cut it down from weeks.
The point with DRAID is that the spare capacity is distributed across all the disks in the vdev, hence the name. Thus all disks are read from and written to when "replacing" a failed disk with the spare capacity.
The combined throughput of the multiple disks is typically much higher than that of the single disk.
In addition they have an additional constraint which allows DRAID to also support sequential resilver.
At least that's my understanding.
There's been several presentations of DRAID on the OpenZFS channel, but the latest one is here[1]. There's also a writeup here[2], though not sure if it's 100% up to date, though should be close enough for getting the concept.
[1]: https://www.youtube.com/watch?v=jdXOtEF6Fh0
[2]: https://openzfs.github.io/openzfs-docs/Basic%20Concepts/dRAI...
Additional speed on top of that is nice of course, if you're using such wide pools.
Not sure how much that matters in practice, I've just got a home NAS with 6 disks to play with.
Is it the fact that you get 12tb of essentially non sequencial io mixed with random reads during the rebuild? Still 20 mb/s seems 1990s slow writing tio a mechanical HD?
Ed: ok, ok, from tfa:
> For 14 wide RAID-Z2 vdevs using 12TB spindles, rebuilds can take weeks. Resilver I/O activity is deprioritized when the system has not been idle for a minimum period. Full zpools get fragmented and require additional I/O’s to recalculate data during reslivering.
Obviously writing 12 tb at 0 mb/s will take time..
If you insert a cold spare to replace a disk — yes, you have to write the full disk.
If you migrate a disk from a hot spare to an active disk, it’s already largely populated, so you only have to write the missing chunk.
At least, that’s my limited understanding/reading. I’m not sure I fully grok what is happening as there aren’t necessarily dedicated spares in this setup.
But I think the difference in your scenario is hot vs cold spares.
When you do a rebuild with a real, distinct physical disk, there's no special benefit, you're still almost certainly going to bottleneck on the write speed of that single disk.
With raid 6 for example, I pop in a new drive, and all the write/repair hits only that drive.
With a virtual spare, all drives end up with massive writes, meaning I might hit write errors, thus borking a raid which would have been ok for reads.
I realise there are 100 'what ifs' and scenarios here, but when my raids tend to drop a drive, it's on heavy write ops, not reads. I'd say 90% of the time.
I imagine ssd cascading/eol write failures might make this virtual spare failure scenario even more likely.
(yes I am aware of sector reallocation delays on spinning disks, I am referring to real EOL scenarios on drives during writes, being far more common for me than reads)
I suppose it works out as follows. With traditional raid, you are rebuilding while degraded from a safety perspective. So obviously, if your configuration guarantees losing `n` drives while healthy does not cause data loss, you could lose data during rebuild if `n` drives failed.
With the distributed hot spare approach, you are no longer degraded from a safety perspective as soon as the distributed spare's data is written. This is faster than traditional raid, as the writes are distributed among many disks. So time-wise you get out of data safety degraded state faster. (It does seem like it might be riskier though, since you are doing heavy writes to many drives while safety degraded, instead of only writes to the spare).
So now you replace the failed drive, and the array rebuilds to make this new spare's capacity distributed. This will take a long time, since you will be writing to basically this whole drive as part of distributing its spare space. (Plus probably shuffling chunks around between many other drives). This could possibly take even longer than traditional raid rebuild. However during this whole process, your array is not in degraded state from a data safety point of view.
Or to summarize: during the long single-drive-write-limited part of replacing the failed drive with a new hot spare, you are not degraded from a safety of data point of view.
Yes. With the caveat that you'll be *reading* from all of the disks to satisfy those writes. In a traditional setup, you'd have multiple striped raidz{,2,3} vdevs. When you swap in a spare, you're only pulling reads from the disks in the same vdev. This will drive the performance from this vdev to the ground. In this draidz setup, you'll be pulling from all of the disks, so the entire pool is stressed at the same rate.
I think the way I'm starting to think about it is that each drive has X amount of spare capacity or buckets. But they are otherwise used. When a drive goes out, the primary data from that drive is distributed across the array to the already-in-place spare capacity. So, yes, after the primary data is mirrored out, the pool is no longer degraded. But, I think the point being made here is that you'll still want to replace the drive, which will still cause *that* drive to be fully written.
I don't think from a data safety or performance perspective pre dRAID vdev really differ too terribly much from raid 50 setups. (They do differ greatly in administrative flexibility, versus the insanely rigid approachs of traditional hardware raid).
But yeah, after replacing the failed drive, ideally dRAID would basically do just what it did after drive failure in reverse. (End result: No longer having any drives holding data for a different vdev than they belong to). Obviously this means writing basically that whole disk, and reading from potentially every disk. But as you mention you are spreading the load out across the whole array, hopefully reducing the overall impact, and being less likely to stress any drive into dying.
Let's say you've got an array with 10 disks and 100TB of data. You experience a single-disk failure.
If you're running a RaidZ, you have to write the lost data (100TB / 10 disks) to the single replacement disk. Writing 10TB of data to a single disk takes quite a while, especially on inexpensive SATA disks.
If you were running a dRAID with equivalent redundancy settings, you have to write (100 / 10)TB to the 9 remaining disks in the array. Each disk will have to write about 1.2 TB of data in parallel.
Writing 1.2TB to each disk should be roughly ~5-10x faster than writing 10TB sequentially.
https://github.com/openzfs/zfs/pull/8853
It's mentioned in the pull request comments. Maybe it will help?
(cf. slide 11 of https://docs.google.com/presentation/d/1uo0nBfY84HIhEqGWEx-T... )
It expires in nine months, which perhaps explains the timing.
The essence of draid is that, instead of keeping a spare drive unused in case one of the working drives fail, it incorporates the spare drive to the array and uses it, but one drive worth of free space is reserved randomly across the entire array.
That way, if one disk fails, the reserved space is used to write the data necessary to keep the array consistent. Because the free space is distributed randomly across the array, the write performance of a single drive doesn't become a bottleneck.
For example I can have a hot spare with a dRAID setup as part of my raidz2 pool?