ZFS on Linux v0.7.0 released
list.zfsonlinux.org
list.zfsonlinux.org
Recently I tried running a simple BTRFS SSD Mirror for only a single KVM virtual machine, I thought compression would be neat since they are just two cheap 120GB SSDs and I wanted to have a spare if one of them gave up the ghost.
At first everything was great but after I ran a btrfs scrub, and it found ~4000 uncorrectable errors, my VM was broken beyond repair and had to be restored from backup. The SSDs SMART data was fine, there weren't any loose cables, everything worked great, no reported (ECC) memory errors... I have no proof but it seems that BTRFS just decided to destroy itself.
I have since moved my SSDs to a ZoL mirror and (after running scrubs every two days for two weeks) had no further corruption, silent or otherwise. To me, this means that btrfs just isn't stable enough for production use - while ZoL is.
This is what eliminates it from home use for me.
It means to expand the pool the only way is to copy everything off (so you need enough spare storage to hold your entire pool), rebuild, then copy it back. The alternative is don't expand, and just add additional pools, but that starts losing benefits quickly (not having to think about where there's free space when adding something, not having to look in multiple locations to find something).
Can anyone who has been using ZFS long-term at home comment? How do you add more space?
The problem is having a 6x3TB pool and turning it into a 7x3TB pool, which is arguably much much cheaper than buying 6x4TB.
But yeah, sizing your pool very generously when first building it is a good idea. I found that 6x 4TB drives in a raidz2 pool was within my budget, and it will take me a long time to fill 16TB.
On BTRFS you simply use the new space or rebalance the pool to use the new disk properly.
On SnapRAID the next scan will add the disk to the parity drive contents.
For low-cost home usage, it is much much more cost effective to only buy single disks and start with small pools over buying large pools or even replacing entire pools.
Remaking the entire pool is also a hassle and incurs unnecessary downtime.
Additionally, not all data is backed up, which I will loose, as this is not important data, it's okay to loose during a house fire, but not just for resizing the disk.
Lastly, this operation would likely take a long time, days probably, I'd rather just be able to just ram in another disk and be done with it.
zpool attach <poolname> <first existing small vdev> <first larger new vdev>
zpool attach <poolname> <second existing small vdev> <second larger new vdev>
[...wait for the resilvering of the new vdevs...] You now have a 4-way mirror with 2 small and 2 large vdevs. Detach the small ones:
zpool detach <poolname> <first old vdev>
zpool detach <poolname> <second old vdev>
Now you have a pool made only of large vdevs. You just give the pool the permission to expand and occupy all this new space:
zpool set autoexpand=on <poolname>
Done. Did it oodles of times on Solaris with SAN storage, but did it at home too, and in the weirdest ways (no SATA ports available? No problem, attach the new drive via USB3, then when finished take it out of the enclosure and install it inside displacing the old drive), sometimes even rather unsafely (creating a pre-broken raidz with 4 good disks plus a sparse file, migrating data off a different pool, then decommissioning that old pool and using one of its disk to replace the sparse file).
If you want, though, you can add capacity by attaching a second raidz2 to an existing pool - this forces you to add 4 new drives instead of one, but it works:
zpool add <poolname> raidz2 <first new vdev> <second new vdev> <third new vdev> <fourth new vdev>
You now have a concat of 2 raidz2's, in a single pool.
Nowhere as elegant as merging a new slice in an existing z2, I concur. Does the job though. Oh and if you are insane you could probably even add a single device, resulting in a concat of a z2 + unmirrored single device. I don't think it will stop you from doing that.
Nope.
It will result in very uneven utilization which leads to poor performance. Not recommended. Is also extremely wasteful. The whole point of the sought feature is to minimize waste and cost. So no, I'd strongly disagree that it does the job.
That said, even if that function existed, expecting it to magically rebalance the whole pool to take into account the existing data is rather unrealistic (it's basically a full rebuild in that case, the most terrifying operation you can perform on any pool/array already containing data).
While I agree with the performance aspects (your data still comes from either the first volume or the second, so you're only using the performance of half your spindles - unless you simultaneously access files in either half), I wouldn't call it wasteful: You are still losing the same % of your total capacity. e.g.:
6x2TB RaidZ2 = 8TB usable (4TB lost to parity)
let's add a second Z2 6x6TB volume:
6x6TB RaidZ2 = 24TB usable (12 TB lost to parity)
Total: 32 TB usable, 16 lost to parity.
What if we had built it from scratch using 6x8TB drives?
Total: 32 TB usable, 16 lost to parity.
Same.
If you build a RAIDZ2 wiht a 3x8TB and a 3x8TB pool, you loose 2+2x8TB = 32TB and have only 16TB usuable.
6x2+6x6 is a mixed pool so it's hard to compare to actual disks.
An alternative, better, comparison, would be a 6x8TB pool and another 6x8TB pool.
This time you loose 32TB to parity and have 8x8=48TB usuable. If you rebuild it to a 12x8TB pool, you loose 16TB to parity and have 80TB usuable.
That is exactly what is required (well, you re-balance the vdev, not the pool) and is what is proposed and has been in the works for basically forever. SUN never prioritized this because it has few uses within the enterprise (but extremely valuable for home users).
Your example is quite misleading. Here's is a much more common and realistic scenario.
6x2TB RaidZ2 = 8 usable, 4 lost to parity.
Now, imagine that you want to expand the pool with 4 TB. Being able to add to a raidz vdev would result in this:
8x2TB RaidZ2 = 12 Usable, 4 lost to parity. (ridiculously cheap, no further waste at all)
Current scenario:
6x2TB RaidZ2 + 4x2TB RaidZ2 = 12 Usable, 8 lost to parity. (expensive, lower performance, more power, more noise, requires more harddrive slots (this is often a very important aspect for home-setups))
The waste is outrageous. And the consequences from this waste reflects every aspect of designing a home-setup (as to avoid the above scenario) with ZFS and is vastly different to how you would design a system with, for instance, raid6 with expansion in mind.
For example, you can start with 1 mirrored vdev, then add another mirrored vdev. You can upgrade vdevs separately, so you'll need to replace just 2 disks to grow your pool.
The only thing you have to keep in mind is that data is striped across vdevs only when written, so if you add another vdev, you won't get performance gains for data which was written just to one vdev.
A mirror can resilver much faster, and it doesn't significantly affect the vdev performance while doing so. A RAIDZ resilver after a disc replacement can take a significant amount of time, and degrade performance seriously as it thrashes every single disc in the vdev.
Allan Jude and Michael Lucas' books on ZFS have tables describing the tradeoffs of the different possible vdev layouts, and they are worth a read for anyone setting up ZFS storage.
A 6x8TB RAID6 has a 0.0002% chance of a URE.
(Assuming URE rate of 10^-14, in reality this rate is lower)
It works pretty well - by the time I'm buying disks that are 4x as large I don't mind throwing the old disks away. I've definitely avoided data loss in scenarios where I'd previously lost data under linux md (which lacks checksums and handles disks with isolated UREs very poorly).
My solution is to use RAID1 exclusively. That means that I can keep attaching pairs of devices if I run out of diskspace. I can never get them out again, however :-)
Once I get it, I hot swap a drive, and start searching for sales on more drives.
Rinse, repeat, all 5 disks swapped, more space appears \o/
When another RAID device is added to the pool instead, it is used as a stripe. Different types of RAID can be combined in this way, for example one can add a mirrored stripe to an already existing RAID-Z. Whether this is desired or not is a different discussion, since the point here is that both scenarios are possible.
Unlike other volume managers, ZFS expands on the disk boundary, rather than physical or logical elements; it's a larger boundary, but since no arcane or complex procedures or knowledge are required and drives are cheap, it's very practical, as well as elegant.
Not sure why ZFS would not auto resize for you. It's the reason I have been using it for several years as my home NAS under a linux server. In my case I just replace all the disk on a RAIDZ with bigger ones and automatically resizes up.
Not that my 10 TB RAIDZ is running out of space any time soon, but as soon as Btrfs gets their shit together I'm switching. At the rate of a few blu rays a year it'll fill up sooner or later.
What's the benefit of RAIDZ over say, you choose to have X copies distributed over your disk(s)?
Answer: zfs: copies=n is not a substitute for device redundancy! source: http://jrs-s.net/2016/05/02/zfs-copies-equals-n/
Here's a discussion about it: https://www.reddit.com/r/DataHoarder/comments/4hbn8v/raidz_v...
Anyone who wants to have real security, relies on off-site backups, isn't that right? And aren't RAID(Z)s slow to recover also? (serious questions, I'm a zfs noob)
RAIDZ: I don't know what stripe-set configuration is good for me and don't want to waste time comparing RAID controllers or if I even need one. Then configuring the beast on the hardware and software side just seems to be too tedious. Why not just (de)/attach another disk and let zfs expand/shrink my total disk space without loosing consistency?
Startup idea: Someone clever should find a flexible storage solution that uses aufs, unionfs etc. to give you the flexibility we need.
### Performance
* ARC Buffer Data (ABD) - Allocates ARC data buffers using scatter lists of pages instead of virtual memory. This approach minimizes fragmentation on the system allowing for a more efficient use of memory. The reduced demand for virtual memory also improves stability and performance on 32-bit architectures.
* Compressed ARC - Cached file data is compressed by default in memory and uncompressed on demand. This allows for an larger effective cache which improves overall performance.
* Vectorized RAIDZ - Hardware optimized RAIDZ which reduces CPU usage. Supported SIMD instructions: sse2, ssse3, avx2, avx512f, and avx512bw, neon, neonx2
* Vectorized checksums - Hardware optimized Fletcher-4 checksums which reduce CPU usage. Supported SIMD instructions: sse2, ssse3, avx2, avx512f, neon
* GZIP compression offloading - Hardware optimized GZIP compression offloading with QAT accelerator.
* Metadata performance - Overall improved metadata performance. Optimizations include a multi-threaded allocator, batched quota updates, improved prefetching, and streamlined call paths.
* Faster RAIDZ resilver - When resilvering RAIDZ intelligently skips sections of the device which don't need to be rebuilt.
What?
* Vectorized RAIDZ - Hardware optimized RAIDZ which reduces CPU usage. Supported SIMD instructions: sse2, ssse3, avx2, avx512f, and avx512bw, neon, neonx2
* Vectorized checksums - Hardware optimized Fletcher-4 checksums which reduce CPU usage. Supported SIMD instructions: sse2, ssse3, avx2, avx512f, neon
* GZIP compression offloading - Hardware optimized GZIP compression offloading with QAT accelerator.
All use vector and/or SIMD instructions on the CPU, thoughh QuickAssist (QAT above) is available on dedicated add-in card also.
Warning to all that use ZFS to host the /boot file system, however: GRUB doesn't presently support all the features added to this release. Be careful about the features you enable!
https://github.com/zfsonlinux/zfs/pull/5769
Judging by the comments from today though it sounds like it will be merged very shortly after this release. So that's good news :D
I'm also wanting to look at this "OPAL" native drive encryption stuff, which requires among other things a UEFI plugin and some other bits. That would be a useful solution also though much more versatile here in ZFS for example doing per-user stuff, and also encrypted send without decrypting
Lots of good changes in 0.7 otherwise!
Full disclosure, I work with Tom at Datto.
(running release candidate on laptop via Arch without issues and current stable on many servers)
Most of the companies here I've never even heard of and the ones I have heard of aren't companies I would take into account when choosing a file system.
http://open-zfs.org/wiki/Companies
In all my years I've never encountered or needed this file system, and every time it's mentioned it sounds like it's more trouble to run it than its worth.
I have more trouble with FreeBSD than ZFS, but even with those troubles ZFS hasn’t let me down.
ZFS is really really amazing, beyond the SAN. It's fantastic for both desktops and servers.
The lack of adaption is problematic, but I use it for all sorts of things.
Like mirroring snapshots between prod and qa, including the database and assets. I can test complicated software upgrades or database migrations quickly and repeatily. This is not something docker can't do by itself.
Much of the stuff I do with ZFS I can now do with btrfs, but ZFS is just so much nicer to work with.
I think that comment has to go on the Best of 'Shit HN Says' corkboard ere long. I mean, when I am faced w/ a technical challenge at work, or at home...why would my first thought NOT be 'What would Google do in this situation?'
ZFS was designed for enterprises. Enterprise software has requirements of high stability, scale etc. Google is probably on the extreme end of these problems. So if Google considered and rejected ZFS it probably means it was not very good.
What probably really happened was that Google was already too invested in Google File System and therefore did not give other file systems a fair chance
Actually, they're big enough that something between what ceph can do and what backblaze does is probably their /archive/ backend.
They probably have all of the fast stuff on 'disposable' temporary copies in RAM or SSDs.
Do you have the expertise that Google has?
Do you have the scale that Google has?
Do you have the problems that Google have?
I had an employee once who, when tasked with designing an API for use by a partner company, copied MS conventions a-la "CreateWindow", when I asked about adding more functionality he said "I'll just do CreateWindowEx" and "IWebBrowser2". His claim was "If MS is doing that and MS is successful, this must be a good way to do it".
After I pointed out that, despite several attempts, there is no MS-Compatible API (this was in 2001, I think reactos and wine had both been started, there had been other attempts at WinAPI emulation, but none was worth anything), and that we had 2-people working on it, not 400, he agreed that perhaps the Unix API is easier to maintain, document and reimplement -- as is evident by many (re)implementations available.
For Microsoft, a labour-intensive-to-maintain-and-labour-intensive-to-document API is a benefit, because they can afford it and they are already holding the dominant market position (making it hard to be compatible with them). For a lean team, maintaing this kind of API is a penalty.
Similarly in this case, what's good for Google and what's good for you are not necessarily equivalent.
Google solves speed, redundancy and reliability by making and managing copies. ZFS makes one single machine the most reliable file system available today. But Google cares not about any single machine.
Do you have 100 copies of every important datum? If you don't, then you shouldn't copy Google.
It's a real shame not to have it more integrated otherwise though in less dedicated applications. It's really hard to go back once you are used to zfs send/recv, transparent compression, etc.