There is a huge performance penalty of course, as all the majority of new data will reside on the latest vdevs - but it generally works.
I do agree this is one of the largest drawbacks of ZFS, but very few filesystems get it right.
There is a huge performance penalty of course, as all the majority of new data will reside on the latest vdevs - but it generally works.
I do agree this is one of the largest drawbacks of ZFS, but very few filesystems get it right.
As long as they're expensive mirror vdevs, right? My impression is that home users want the efficiency of RAID-6 and they want incremental expansion (regardless of whether this combination is "good for them"). ZFS can't do that.
That, of course, is really not a concern for enterprise use cases.
For a 6x8TB and assuming a (optimistic) 10^-16 URE, you get 3.5% failure rate for a RAID5 array, 0.7% for a RAID10 array and a 1.06e-08% failure rate for RAID 6.
Why be greedy for all that performance? Most home-grade NAS or even some business-grade NAS isn't used for performance sensitive operations, more like Word Documents and Family Pictures, stuff you don't want to loose.
I'd rather take safety over performance here.
Regardless, the determining factor here is how much data do you need to read in the case of a failure to rebuild the array. RAID1 wins every single time because you cannot read less than the single drive you need to replace.
However, the chance of failure for a RAID6 of 100x10TB disks is less than 0.482% after 1'000'000 rebuilds.
RAID1 is space-inefficient, a 100x10TB RAID array might never fail but it has only 10TB of storage space.
A RAID10 Array has a 14% failure chance for just 4x2TB disks using 10^14 failure rates.
RAID1 and RAID10 are definitely not the way forward, it is less secure, something that should be immediately apparent if you read the link in my previous comment.
A 10 Disk RAID6 with 10TB disks is more reliably than a RAID10 by multiple orders of magnitude and more space efficient than a simple RAID1.
In the case of a RAID1, noone uses 100 mirrored drives. You use RAID10, and in the case of a failed disk, you must read 10TB to recover. With the same URE, we'd see on average 8 URE for every 10 rebuilds, or around 2 orders of magnitude less failure rate compared to the RAID6 example.
During a RAID6 rebuild, a URE is non-critical as the Array can recover the data with one lost disk an a URE on any other disk during the stripe rebuild.
The only critical error would be a URE on two disk on the same stripe, 80 URE's during a 990TB rebuild have an amazingly low chance of having two UREs on the same stripe on two seperate disks.
In case of the RAID10, you get 8 URE's over 10 rebuilds, which aren't recoverable unless you have 3 disks. So you'll corrupt data.
edit: URE of 10^14 is what most vendors specify for consumer harddrives, 10^16 is closer to what people encounter in the real world but 10^14 is considered the worst case URE rate.
URE does not have to corrupt data, if you use a proper filesystem with checksumming such as the ZFS.
When a disk fails, a RAID10 is simply in a far better position as it only have to read a single disk, and it doesn't have any complicated striping to worry about. Just clone a disk.
No but afaik there is no way to recover data once ZFS has declared it corrupted. (ie, no parity)
>The strain of rebuild have been known to kill many arrays, both RAID5 and 6.
I haven't actually encountered that yet. Despite that, a RAID 6 can loose a disk, so as long as you don't encounter further URE's after loosing another disk, it's fine.
If you're worried about that, go for RAIDZ3 or equivalent. With something like SnapRAID you can even have a RAIDZ6, loosing 6 disks without loosing data. The chances of that happening are relatively low.
>When a disk fails, a RAID10 is simply in a far better position as it only have to read a single disk
A RAID 10 is in no position to recover from URE's once a disk has failed unless you reduce your space efficiency to 33%.
I personally favor not corrupting data over rebuild speeds.
Striping might be complicated but that doesn't make it worse.
It might be acceptable too loose a music file, but once the family image collection gets corrupted or even lost on ZFS because a disk in a RAID 1 encountered a URE, it's personal.
I'd rather life with the thought that even if a disk has a URE, the others can cover for it. Even during a rebuild.
I'd love to run ZFS, but I can't because I need to be able to add more drives as I buy them.
You could also add another raidz vdev if you have the space in the server, eventually they should level out depending on workload.
"Best" solution, read most expensive, would be to copy the data over to a new server you've configured for the new load/capacity.
Both have high investment costs for something a plain mdRAID, snapRAID or LVM RAID can achieve far simpler with better results.
A 4TB Seagate IronWolf drive costs $129 off Amazon, buying two to add as a new mirrored vdev to my TrueNAS box isn't outrageous.
My own personal budget is very limited, buying two IronWolf HDDs in this case is easily a good chunk of my monthly income. If I can build a NAS that can expand with single drives as needed, it's more cost effective for me.
And I imagine a lot of others have the same problem.
In the end, by buying 2 drives when a single drive could have solved the problem equally well: expanding your space by 4TB, is wasting money. Period.
I actually went custom because those off-the-shelf boxes are either very expensive or have weak CPUs so can't be used for video transcoding very well - it's cheaper to build it yourself, much cheaper if you already have old hardware to dedicate to the task.
Having an L2ARC helps out quite a bit, but only having 32GB of memory and wanting to keep most of it for the L1ARC means I still hit my spinning disks regularly (and mirrored vdev's help read IOPS tremendously in this case).
But it doesn't look like there's been any movement on an implementation and it seems like it's high effort and mostly home users who want this, not enterprises who might be willing to pay for.
Ah well, guess I'll stick with mdraid for now.