ZFS Gotchas
nex7.blogspot.ch
nex7.blogspot.ch
Other file systems and RAID setups do not checksum your data. If you have a mirror of two disks (RAID1) and during a read two blocks differ from one another, most (all?) RAID controllers (hardware and software) will simply choose the lower numbered disk as canonical and silently "repair" the bad block. This leads to loads of silent data corruption on what we might consider a reliable storage solution.
By contrast, ZFS will store both the block and its checksum and will use the block with the correct checksum as canonical.
In other words, if you care about your data at all, use ZFS. I am frankly surprised it's not the standard file system for most situations as it is the only production filesystem that can actually be trusted with data.
P.S.: I have been told that at least on Linux if you have more than two drives, the Linux software RAID controller will try to choose the version of the block that is agreed upon by most drives, if that's possible. This is no guarantee, but it's better than randomly choosing one version.
P.P.S: BTRFS and friends seem to not yet be as production ready as ZFS. Conversely, ZFS works beautifully on Linux thanks to the ZFS on Linux project.
I'll gladly take silent data corruption over a hard total failure of the entire volume any day. At least I have a chance at fixing the former, and for my use cases that's preferable.
Every time ZFS on linux is brought up it's always met with skepticism, just as a lot of people trying out ZFS in virtual machines discovered the hard way (lost volumes) there are a lot that can go wrong (and it does seem like it can go wrong) and any version running on linux will not have been tested as much making it a risk. And that's the thing you are trying to minimize with using ZFS on Linux in the first place...
Now BTRFS isn't ready for production yet so that leaves, well, nothing if you want checksums and cheap snapshots.
I guess I'll have to wait another 5 years and see if BTRFS is ready, until then it seems the only thing I can hope for is luck.
I'm planning to build a (Linux) development workstation soon and was intending to run ZFS on at least a proportion of the disks in the system, if not the whole thing. But building a workstation that supports ECC appears to be ridiculously hard and expensive - most of the Haswell CPUs don't support ECC. Can anyone point me at a sensible CPU/motherboard combo that's comparable to something like an i7-4770k/Gigabyte GA-Z87MX-D3H?
(As an indication of the state of the component market, pcpartpicker.com doesn't support searching CPUs or motherboards by ECC support, but it does allow you to search by what colour the motherboard is!)
By comparison, I've had issues with virtually every other filesystem that exists for Linux: ext(2,3,tho not 4), reiserfs, and especially the much hated xfs, death to xfs etc :)
I run his hardware, picked because... well cheap:
GIGABYTE GA-990FXA-UD5 ( AMD 990FX/SB950 - Socket AM3+ - FSB5200 )
4 x DDR3 4GB DDR1866 (PC3-15000) - KINGSTON HyperX [KHX1866C9D3/4G]
or
4 x 16gb of the same type as above
AMD FX 8-Core FX-8320 [Socket AM3+ - 1000Kb - 3.5 GHz - 32nm ]
The ZFS all include root + swap running on ZFS with FDE, this sort of thing for workstations: NAME USED AVAIL REFER MOUNTPOINT
rpool 13.2G 25.9G 30K none
rpool/root 11.1G 25.9G 5.10G /
rpool/root/home-users 6.02G 25.9G 6.02G /home
rpool/swap 2.06G 27.4G 646M -
servers: NAME USED AVAIL REFER MOUNTPOINT
rpool0/ROOT/ROOT 6G 1.42T 198K /
rpool0/ROOT/swap 32G 1.42T 700M -
rpool0/ROOT/BACKUP 676G 1.42T 676G /nfs/backup
rpool0/ROOT/HOME 14.1G 1.42T 14.1G /home
rpool0/ROOT/MUSIC 358G 1.42T 358G /nfs/music
rpool0/ROOT/portage 14.9G 1.42T 14.9G /nfs/portage
rpool0/ROOT/vmware 45.9G 1.42T 45.9G /vmware
The server hosts several virtual machines, many of which have their own root over NFS, the point being that the ZFS gets plenty of use.Basically I wanted something that doesn't intertwine the concept of file-system with the concept of disks or partitions, but as space. So changes to the underlying pool of disks (or partitions, or loopback files, etc) don't mean so much painful rebuilding for me, something I spent many hundreds of hours on in the past, often at stupid o'clock in the morning. Did I mention how much I hate XFS? That's why :)
Basically you can roll the dice with AMD, or just go with Intel (what I'll be doing). If you want reasonable performance, get a socket 1150 xeon v3 (basically a re-branded i7 with ecc support). The low-mid xeon v3 appear to score pretty well on a price/performance ratio. Can't be overclocked, and you'd probably get more single thread performance/price with an i5 -- but apart from that they seem pretty solid. As others have mentioned, (some) i3s are also an option -- but as you'll likely have to pay a little more for a main board with ecc support, that doesn't make much sense as I see it.
The final option is to get one of the newish atom based Avoton boards, like ASROCK C2750D4 INTEL C2750 AVOTON OCTACORE MITX -- not for a workstation, though. But should be nice as a NAS.
If anyone knows of a useful overview of AMD cpus and mainboard combinations that support ECC -- please let us know. I've yet to find anything beyond the anecdotal "Here's my build, and it's got ECC sticks in it, and maybe the ECC is actually enabled, but I haven't really checked."
You can still get most of the benefits on a workstation with non-ECC RAM, consider that any other file system would have the same challenge to deal with on non-ECC RAM.
So, as in all things, it's a money/benefits tradeoff: Does the data you're working with warrant spending a ridiculous amount of money?
[1] http://ark.intel.com/search/advanced?s=t&FamilyText=4th%20Ge...
[2] http://www.supermicro.com/products/motherboard/Xeon/C220/X10...
Those boards and their SAS chipsets (LSI SAS2008, etc.) tend to be well supported on platforms capable of running ZFS, like Solaris, FreeBSD etc.
Other typical goodies include IPMI (integrated IP-KVM), good quality dual NICs and generally good quality components.
It has 2 USB 3.0 ports, supports up to 16GiB of RAM, and has 6 SATA 6Gb/s ports, so I think it would also work well for a workstation board, if you added a video/audio card -- HDMI out is the only glaring omission for a workstation.
I'm using it in my FreeNAS fileserver and have been happy with it.
Granted, most application programmers don't deal with the file system in any meaningful way - most interaction is deferred to other processes (e.g. the datastore). But I for one would be interested in certain silly things like, oh, writing a program that can (empirically) tell the difference between a spinning disk and a solid state disk, or what file system it's running on, just from performance characteristics. Other fun things would be to determine just how fast you can write data to disk, and what parameters make this rate faster or slower.
Many of these considerations don't have anything to do ZFS per se, but come up in designing any non-trivial storage system. These include most of the comments about IOPS capacity, the comments about characterizing your workload to understand if it would benefit from separate intent logs devices and SSD read caches, the notes about quality hardware components (like ECC RAM), and most of the notes about pool design, which come largely from the physics of disks and how most RAID-like storage systems use them.
Several of the points are just notes about good things that you could ignore if you want to, but are also easy to understand, like "compression is good." "Snapshots are not backups" is true, but that's missing the point that constant-time snapshots are incredibly useful even if they don't also solve the backup problem.
Many of the caveats are highly configuration specific: hot spares are good choices in many configurations; the 128GB DRAM limit is completely bogus in my experience; and the "zfs destroy" problem has been largely fixed for a long time now.
It was... slow. But the blinky lights were pretty awesome.
There's actually a trick you can use to create a failed ZFS array (e.g. if you want to create a 5+1 array with only five disks), where you use dd to create a sparse file of the appropriate size, which lets you create a '1 TB' file while only writing 1 byte to disk. Add it to your zpool along with the rest of your disks, then fail it out, remove it, and replace it with a physical disk.
It's a good way of taking a drive with 2 TB of data on it and adding it to a ZFS filesystem which includes that disk while also keeping the data, but without using a separate disk (or copying the data twice).
Just remember to keep back ups. It doesn't matter how good you might think your storage array is; always make back ups.
However, none of these diminish what I think is the main use case of ZFS: very robustly protecting against corruption below the total-disk-failure level.
This might kill my dreams of a ZFS NAS...
The critical concept here is that where other systems use physical disks, ZFS uses vdevs. vdevs are 1 or more disks but are presented to the RAID system as a single 'storage entity'. Thus, you don't add disks to a storage pool, you add vdevs.
zpool replace [-f] pool device [new_device]
Replaces old_device with new_device. This is equivalent to attaching
new_device, waiting for it to resilver, and then detaching
old_device.
The size of new_device must be greater than or equal to the minimum
size of all the devices in a mirror or raidz configuration.
new_device is required if the pool is not redundant. If new_device is
not specified, it defaults to old_device. This form of replacement
is useful after an existing disk has failed and has been physically
replaced. In this case, the new disk may have the same /dev path as
the old device, even though it is actually a different disk. ZFS
recognizes this.
-f Forces use of new_device, even if its appears to be in use.
Not all devices can be overridden in this manner.Lets say I have 6 SATA ports. I have 4 drives that I collected from various computers that I now want to unify in a home built NAS:
A. 1TB
B. 2TB
C. 4TB
D. 1TB
E. -empty-
F. -empty-
Now all my drives are full and I want to either add a disk or replace a disk; How do I: 1. replace disk A. with a 4TB disk
2. add a 4TB disk on slot E. 1. connect 4T disk to F, "zpool replace pool A F"
2. connect 4T disk to E, "zpool add pool E"
Where A, E, and F are virtual devices (vdevs) provided by your OS. On my FreeBSD system they are da0, da1, etc.Obviously if you only have 4 drive slots and have to remove a disk to replace it then you're going to have less redundancy while you're in the process of doing the replacement.
You can upgrade the capacity of a pool, but you have to upgrade each drive, one at a time, with time to resilver in between, and the bigger capacity doesn't come online until all drives are incorporated.
You could add another 4x3TB vdev to the pool to add more storage to the pool, but once you did you couldn't remove it.
When it is not exported, the pool stores the device names of its component disks, and if the wrong disks end up on the wrong devices, you get the "corrupted data" problem even though the data really isn't corrupted.
Is there any reason why one might choose that for a small office / home office setup instead of ZFS? Would it make setup/expansion easier.
Taking it in another direction.
I've been following this LeoFS project which basically replicates Amazon's S3 storage (so any S3 clients can work with it).
Its claim is one can add nodes to the cluster to expand its storage capability. Given enough physical space it might be cheaper to find 3-4 machines (towers) with 4 drive bays than building one new server with 12+ bays. Then use linux's fuse client to access it.
At some point 2 older towers or hardware with only 4 drive bays might be cheaper than one motherboard and server tower that takes 8 or 12 drives.
That is why I was wondering about the clustered/distributed FS option.
If you get a full tower, you should be able to fit a couple of (or equivalent): http://www.newegg.com/Product/Product.aspx?Item=N82E16817198...
(3.5" & 5.25" Black Tray-less 5 x 3.5" HDD in 3 x 5.25" Bay SATA Cage) in front -- if yo feel you need front-accessible drives.
I'd say you normally want less boxes to take care of (even if failure will be more catastrophic) on such a small scale. Depends on what you need of course.
http://arstechnica.com/information-technology/2014/01/bitrot...
Unrelated: I really don't think dang's post here warranted downvoting. I expect that it's being done by folks who really want HN's moderators to know they're unhappy about titles being changed as capriciously as they sometimes are. This post is an actual moderator soliciting actual input on how to make at least one instance of that problem better, however, and I think that should be encouraged, not dinged.
It's ok for dang to get dinged. But "capricious"? No. I realize it sometimes seems that way to people paying sporadic attention, but every change we make is in keeping with the HN guidelines, and those are hardly secret or unclear.
We make mistakes, of course, and are always open to improvement. The helpful way to criticize an HN title is to suggest a better one.
Better policy a): Don't change the titles users submit. Trust the community to flag linkbait titles.
Better policy b): Don't allow users to submit a title. Automatically scrape the title from the linked page. Have moderators change the titles when they're unhelpful or misleading (this is much less likely to annoy users than the current system, because the moderator wouldn't be replacing the submitter's title, it would be the HN system replacing another part of the HN system).
edit: Suggest a better title rather than suggest a better policy? How is one supposed to suggest a better title? There is no form for that, and it's not something to clutter the comments with.
So by changing it to a more boring title, you ensure that fewer people read the article.
Brilliant.
Incidentally, "Things Nobody Told You About ZFS" is the subtitle of the article in question, and arguably is a more informative description of its contents than "ZFS: Read Me First".
Oops, I'm doing that on my home NAS. Does anyone know why this is bad?
Now, the error detection schemes at the disk level may be insufficient. I don't know enough about how it's done on modern drives (but I suspect that every manufacturer has its own scheme).
Hard drive capacities have been rising much faster than the unrecoverable read error rate has been lowering.
If all ports on the controller were in use, and you yanked c4t7d0 and slid in a new drive, it would become c4t8d0 (unless you used lsiutil to remove the persistent mapping).
Other HBAs may have similar features.