ZFS Is the Best Filesystem For Now
blog.fosketts.net
blog.fosketts.net
How is L2ARC not "true hybrid"?
> And no one is talking about NVMe even though it’s everywhere in performance PC’s.
Why should a filesystem care about NVMe? It's a different layer. ZFS generally doesn't care if it's IDE, SATA, NVMe or a microSD card.
> can be a pain to use (except in FreeBSD, Solaris, and purpose-built appliances)
I think it's just a package install away on many Linux distros? Also installable on macOS — I had a ZFS USB disk I shared between Mac and FreeBSD.
Also it's interesting that these two sentences appear in the same article:
> best level of data protection in a small office/home office (SOHO) environment.
> It’s laughable that the ZFS documentation obsesses over a few GB of SLC flash when multi-TB 3D NAND drives are on the market
Who has enough money to get a mutli-TB SSD for SOHO?!
It doesn't persist across import/export or reboot.
It is demand-filled.
Not all data in the main vdevs are eligible for l2arc.
There is memory overhead for l2arc buffers.
There is CPU overhead in processing l2arc headers.
Once a buffer is in l2arc it stays in l2arc until the underlying data is overwritten or destroyed, or until the l2arc has filled up and the buffer is replaced with fresher data.
A true hybrid in the zfs context would let one pin a dataset or zvol onto a particular vdev, or pin only the (zfs) metadata (or a subset thereof) of a dataset, zvol or pool to a particular vdev.
Openzfs will eventually get both persistence and this form of true hybrid.
Unfortunately automatic migration by zfs of hot data to low-latency vdevs and cool data from low-latency vdevs is not really possible without solving the infamous block-pointer-rewrite problem.
Because at some point the filesystem becomes a bottleneck. ZFS was designed with the assumption that CPUs would be way faster than storage. When you get speeds over 10GB/sec, [0] you are going to spend a lot of time checksumming all that data.
[0] http://www.seagate.com/ca/en/about-seagate/news/seagate-demo...
https://github.com/zfsonlinux/zfs/issues/4789#issuecomment-2...
https://cloud.githubusercontent.com/assets/472018/13333262/f...
Update: At least 1 person didn't get this joke :D
What else are you going to use the two CPUs in a filer for?
ZFS is well pipelined for multicore throughput.
L2ARC is only ever a cache and although it can be on SSD, ZFS generally won't use much more than a few tens of GB. ZIL isn't even a cache and is only used for synchronous writes. Check any ZFS tuning guide and the gist will be "just buy more RAM or create an all-SSD pool" rather than trying to wedge an SSD into L2ARC or ZIL.
We have a TrueNAS appliance with a 480GB L2ARC and a small 120GB ZIL (never fully used) used for image/document storage by our ECM suite. Our metadata usage on the filesystem is astronomical due to having billions of small (<16KB) files, the L2ARC may not do a whole lot for actual data (since most of it isn't hit frequently enough to be eligible to be stored in the ARC/L2ARC) but it's instrumental in maintaining performance with such massive amounts of filesystem metadata.
Pretty much every enterprise storage solution designed today is hybrid (SSD plus HDD) or all-flash and includes lots of advanced availability features.
By this do you mean simply that there is both SSD and disk storage around, or do you mean that the storage system transparently chooses where to store particular things without the apps having to care.
Because the latter thing is not my experience.
These systems will transparently "promote" hot blocks to flash or faster spinning disk, etc and demote cold blocks based in usage patterns and policy, without the app being aware.
There are a lot of different ways to do it, ranging from tiering within the array to a storage virtualization solution that can tier data across different storage platforms. In one case I worked on a project where that virtualization tech was used to consolidate 10 data centers to 1, pretty much transparent to the end users and mostly transparent to the folks running apps. (Exceptions were ostly apps that built their own Storage HA)
L2ARC in my understanding is only for reads, whereas ZIL is the write ahead log. Ideally, ZFS would "combine" the two into a MRU "write through cache" such that data is written to the SSD first, then asynchronously written to the disk after (ZIL does this already) but then, when the data is read back, it's read back from the SSD.
Did the ZIL and L2ARC concepts come up before SSD was widely available? Especially the ZIL seems very much optimized for crazy enterprise 15k rpm spinning rust. Memory and SSD access characteristics are so different from spinning disks; I don't know why ZFS separates ZIL and L2ARC.
By default, the ZIL is written on the same disks as where the data will be stored, but an external device (aka SLOG) can be added. From that point on, your write IOPS will be limited by this SLOG device, so normally you add a more expensive fast disk as SLOG device to increase your write IOPS.
Yes. After they realized that their initial claims about not needing such things was bullshit (which some of us had told them at the time) but before SSDs became common.
It's easy to install if you want to use it as an additional filesystem. But if you want to install e.g. RHEL on root ZFS, it's quite an adventure. Even Ubuntu with first-class support for ZFS does not support ZFS on root from the box. Actually I don't understand it. Of all the features, snapshots looks like killer feature for Linux distributions. Make snapshot before upgrade, allow easy rollback if upgrade gone wrong. Something like Windows restore points, but much more reliable.
But, the tooling is still rather new and untested. At work we are still stamping out upstream bugs that really shouldn't exist, but the ecosystem is certainly getting better. It's only been a couple years since the ZFS on Linux project has gotten remotely any uptake by the distro folks, and even now that interest is rather tepid due to the licensing issues.
If ZFS was GPL I think it would have been the default filesystem on Linux for quite some time now.
Which is totally fine for a NAS! But yeah, I like my ZFS root on all the FreeBSD installs :)
> snapshots looks like killer feature for Linux distributions
True. FreeBSD and of course Solaris/illumos had boot environments for a long time, it's an excellent feature.
The need for a snapshot to be on disk when sending, what you cite says you get an obscure error "stale NFS file handle" not that you get missing files. If you're silently getting missing files on a send/receive, that's a bug and should be reported.
I use snapshots quite a bit both for data, backup/replication with send/receive, and for root fs, and haven't had problems with it. The known problem with Btrfs snapshots is that they are deceptively cheap to create, but become expensive later on to delete due to back reference searching, freeing extents (or not if they're still held by other snapshots) and updating metadata. Dozens to small hundreds aren't normally a problem, in my use case I don't notice performance problems.
This is true with any Linux-supported FS on most modern PCs since UEFI can only read VFAT partitions.
https://www.amazon.com/Crucial-MX300-Internal-Solid-State/dp...
I was contemplating a build with 2 of these in a RAID 1 configuration for my next homelab server.
Personally I run a gaming (windows) desktop at home and an always-on UPS-backed homelab server that handles minor ops tasks (mostly backing up side projects and some ETL) + provides dev VMs.
My home office "budget" is ~$2k/year. My gaming/work desktop is ~4 years old, represents ~$2.5k of that budget. Monitors/peripherials/desk/chair generally eats another $2k and are of a similar age. I generally spend ~$2.5k on the dev server. I then usually toss ~$1k into a laptop.
I could easily see someone who purely works from home (rather than 1-2 days a week) operating with a larger budget and genuinely needing a ZFS setup of 2TB SSDs.
Realistically, $2-3k/year is 2-5% of the sort of salaries we see on HN given we make a living at this sort of thing...it isn't surprising people would spend that kind of money to me.
https://hurdlr.com/blog/software-web-developer-tax-deduction...
Keep in mind "business equipment" certainly qualifies for such a dev server so you won't be paying taxes on it if you itemize as well.
Even factoring in my homelab spend I don't get any more itemizing than taking the standard deduction as a married individual making $85K, and I spent well close to $2000 on it last year.
I think what that may be referring to is that the ARC is in-RAM and obviously cleared on a reboot, so as a result L2ARC on an SSD is also not persistent. After a reboot, you have to allow the ARC to fill up, then as it evicts data from the L1ARC it's pushed to L2ARC. Until that happens the SSD is not used at all.
Whatever.
What would be really nice is raw SSD/storage access so that ZFS (or other FS) could manage all the wear leveling and bad block mappings.
It absolutely belongs in the filesystem. When you're doing RAID of any type across the devices, you need that layer to manage the underlying media. A single device view will never appropriately manage wear leveling and garbage collection.
There's a reason companies like NetApp have been working with drive vendors to have more control over the underlying media:
http://www.samsung.com/us/labs/pdfs/2016-08-fms-multi-stream...
With this in mind, I don't see why the disk cannot handle this in the firmware. As long as there is enough free NAND on the drive, it can manage wear level and GC just fine assuming that it gets TRIM commands.
ZFS includes volume management, _naturally_. I say naturally, but that was a radical idea 15 years ago. Even now it's not universally accepted, but it's quite correct!
The author of the article seems to assume we should all be trusting SSD or "hybrid storage" firmware to properly handle this sort of thing for us like nice black boxes.
I think this was one of the major problems ZFS was designed to solve. To make storage hardware more simple the idea was to move a lot of this logic into the OS (especially as large RAM sizes got cheaper). It's why having a RAID controller sitting under your vdevs is advised against.
You get raw NAND access on like, home routers. OpenWrt/LEDE uses JFFS2 on that.
And like… 2 TB of cache is a bit high for SOHO NAS, and if you go full SSD for storage, you'd want two of them for a mirror and that's 1000 USD already…
Upgrading $100 hard drive to $400 SSD results in like, 500% improvements in storage speed. If you have any storage-related task... such as video editing, handling of large datasets and whatnot... the SSD will have a far bigger impact on your productivity than any CPU upgrade.
And if you want 6 or 7 of those puppies, in a non-diy server through a vendor which offers support, 10GE and with enough ECC and CPU to handle that IO load/throughput, you're in for a treat :)
On the other hand, I don't think many people would be hitting any limits where this matters.
Also, be sure to have 8+ GiB system RAM available at all times or performance is gonna suck.
Having your root on a filesystem that is provided with your kernel is not ideal situation, update issues make your system basically unbootable.
As far as I'm aware, it's not persistent. Reboot, and your cache of recently accessed files is gone.
I understand Intel is segmenting reliability into higher-priced business gear, but as a developer that depends on this stuff for their livelihood the current status quo is not acceptable.
Linux should have better options since profit margins are not an impediment.
Granted, that's still not a laptop CPU. But it's not a Xeon either.
But I remember when ECC memory dropped out of favor. I wish it had never happened.
we will see how their mobile chips stack up
It uses the Intel G4560 which supports ECC, and is very inexpensive. It's low power too.
Not to mention that no one could tell me if you could install ECC memory or if it would be used by the firmware if it even worked at all.
Not low power though, 1.3v each 32G DIMM.
I had to buy the memory myself but I was looking for "better" memory anyway.
I can't find those DIMMs, 16GB DIMMs are the biggest I can find.
I can't find them online.
For reference they're Samsung branded, I can take a photo if you like; you can see the "width" of the channel in linux which tells you if you're using ECC or not.
I can see the full "width" in linux, which indicates that the OS can see it.
I can't change the code since it is third party. The only way I saw to easily fix it was on system startup to copy the fonts under a new subdir in /tmp (so in tmpfs, ie RAM, no ZFS at all there ) and then softlink the dir the product was expecting to the new dir off of /tmp, eliminating the ZFS high-volume multiple-reader bottleneck.
Never had this problem with the latest EXT filesystems on my volume groups on my Linux VMs with the same 3rd party library and same volume of throughput.
https://github.com/zfsonlinux/zfs
From reading the pull requests and issues on that repo I've got the impression that the next release 0.7.0 will be quite step forward and there seems to be quite sophisticasted work to tackle performance issues.
* http://www.open-zfs.org/wiki/Main_Page
Unfortunately the next generation HAMMER2 [1] filesystem's development is moving forward very slowly [2].
Nevertheless, kudos to Matt for his great work.
[0] https://www.dragonflybsd.org/hammer/
[1] https://gitweb.dragonflybsd.org/dragonfly.git/blob_plain/HEA...
[2] https://gitweb.dragonflybsd.org/dragonfly.git/history/HEAD:/...
If you are using a recent Linux Kernel I can suggest you to use dm-integrity [0] (optionally with dm-crypt) with your favorite filesystem. It's not erasure coding but it can help detecting silent data corruption on the disk.
“Not quite finished - it's safe to enable, but there's some work left related to copy GC before we can enable free space accounting based on compressed size: right now, enabling compression won't actually let you store any more data in your filesystem than if the data was uncompressed.”
What is the point of having compression enabled if you can’t store more data than you could if it was uncompressed? Shouldn’t they just say “compression mechanism works but not useful yet. as of now it is just extra overhead”..
Now that Ubuntu has ZFS build-in by default, I'm seriously considering switching back, and since I too have been burned by Btrfs, I guess I'll stay with ZFS for quite some time. Still, the criticism of the blog post is fair, e.g. I was only able to get the RAM usage in control after I set hard lower and upper limits of the ARC as kernel boot parameters (`zfs.zfs_arc_max=1073741824 zfs.zfs_arc_min=536870912`).
[1] https://github.com/zfsonlinux/zfs-auto-snapshot
[2] The coolest feature is the virtual auto mount where you can access the snapshots via the magical `.zfs` directory at the root of your filesystem.
This.
We[1] offer ZFS filesystems in the cloud[2] and one of the nicest things to explain to customers is that they don't have to think about "incrementals" or "versions" or retention in any way. They can just do a "dumb rsync" to us (mirror) and our ZFS snapshots, on their schedule, will do the rest.
In the event of a restore, the customer just browses right into "5 days ago"[3] and sees their entire offsite filesystem as it existed 5 days ago.
[1] rsync.net
[2] http://www.rsync.net/products/platform.html
[3] rsync.net accounts have a .zfs directory
"Once you build a ZFS volume, it’s pretty much fixed for life."
The ease of growing/shrinking existing volumes and adding/removing storage is why I made the decision to go with btrfs when I rebuilt my home file server.If you build a volume, you can add more zdevs to it.
My suggestion is to avoid RAIDZx and go with mirrors. To add more storage, just add another pair of drives and add it to your pool.
Here is what my main `tank` looks like:
pool: tank
state: ONLINE
scan: scrub repaired 0B in 4h59m with 0 errors on Tue Jul 4 12:47:02 2017
config:
NAME STATE READ WRITE CKSUM
tank ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
wwn-0x50014ee20eba1695 ONLINE 0 0 0
wwn-0x50014ee20eba337e ONLINE 0 0 0
mirror-1 ONLINE 0 0 0
wwn-0x50014ee2640efb3a ONLINE 0 0 0
wwn-0x50014ee2b964ef3f ONLINE 0 0 0
mirror-2 ONLINE 0 0 0
wwn-0x50014ee2b964f2d5 ONLINE 0 0 0
wwn-0x50014ee2b964f4f4 ONLINE 0 0 0
cache
wwn-0x5002538d704e9ff1-part4 ONLINE 0 0 0
errors: No known data errors
You can simply add another mirrored pair and get more storage. Or, you can update an existing mirror by adding new drives, then removing the old ones. Mirroring has better performance and IOPs than striped raid too. If you feel that mirrors might be less reliable than RAIDZx, I'd suggest you read up about it - generally mirrors are more reliable (and you could mirror 3 drives if you liked, get more throughput and IOPs - or have a hot spare).The more complex answer is that over time as you write more data, ZFS will re-balance the distribution of data over the available zdevs.
[1] if you're not careful, this is actually very likely to happen. If you purchase 2 disks at the same time, and they are under the same usage patterns (which they will be when working in a pair), under the same temperature then they are very likely to die at the same time.
LVM on the other hand makes it easy to do snapshotting, RAID, volume resizing, adding more drives, etc - at the cost of some performance in some cases. For the most part, it's the best trade off available if expandability is a requirement and you can't get away with simple mdraid.
Regardless, I really haven't had any operational problems with btrfs for my 3 TB or so of data and when it has managed to get wedged because it couldn't allocate more space to the metadata pool during the initial import of data from my old array I fixed it with a simple rebalance command.
A cron set up to run a minor rebalance weekly helps ensure you never run into that situation in practice and I've not lost any data so for now I'm comfortable using btrfs as my primary filesystem.
I am a little concerned about the longevity of btrfs in general though because it hasn't been receiving a lot of development work lately.
?
Doesn't ring a bell.
A lot of problems with Btrfs sounds like hardware problems of one sort or another. If it's a legit bug the only way such things get fixed is to report it to the developers <linux-btrfs@vger.kernel.org> with complete logs and system information. Did you?
The fsck is definitely hit or miss, but my perspective is the emphasis is on fixing bugs that obviate file system problems in the first place. The reality is an offline fsck for large file systems is just not scalable, so the best bang for the buck is bug squashing.
And there's a tons of that happening.
For the initial 4.12 pull (now done, and probably had a few dozen changes during rc's) 40 files changed, 1629 insertions(+), 834 deletions(-) https://lkml.org/lkml/2017/5/9/510
For the initial 4.13 pull 47 files changed, 1707 insertions(+), 1400 deletions(-) https://lkml.org/lkml/2017/7/4/436
It happened to me while on vacation -- I only had satellite internet and had to fix it over ssh at single characters per second.
It might have to do with a bug I saw mentioned on the list that referred to the free space map getting corrupted (but easily fixed with a -oremount,clear_cache). So I also run that monthly as well.
I've been too afraid to remove the cron job even though I moved my CentOS7 to the "mainline" kernel rpms from http://elrepo.org
Also, I only trust RAID1 + Crashplan.
https://btrfs.wiki.kernel.org/index.php?title=RAID56&diff=30...
There is a huge performance penalty of course, as all the majority of new data will reside on the latest vdevs - but it generally works.
I do agree this is one of the largest drawbacks of ZFS, but very few filesystems get it right.
As long as they're expensive mirror vdevs, right? My impression is that home users want the efficiency of RAID-6 and they want incremental expansion (regardless of whether this combination is "good for them"). ZFS can't do that.
For a 6x8TB and assuming a (optimistic) 10^-16 URE, you get 3.5% failure rate for a RAID5 array, 0.7% for a RAID10 array and a 1.06e-08% failure rate for RAID 6.
Why be greedy for all that performance? Most home-grade NAS or even some business-grade NAS isn't used for performance sensitive operations, more like Word Documents and Family Pictures, stuff you don't want to loose.
I'd rather take safety over performance here.
Regardless, the determining factor here is how much data do you need to read in the case of a failure to rebuild the array. RAID1 wins every single time because you cannot read less than the single drive you need to replace.
However, the chance of failure for a RAID6 of 100x10TB disks is less than 0.482% after 1'000'000 rebuilds.
RAID1 is space-inefficient, a 100x10TB RAID array might never fail but it has only 10TB of storage space.
A RAID10 Array has a 14% failure chance for just 4x2TB disks using 10^14 failure rates.
RAID1 and RAID10 are definitely not the way forward, it is less secure, something that should be immediately apparent if you read the link in my previous comment.
A 10 Disk RAID6 with 10TB disks is more reliably than a RAID10 by multiple orders of magnitude and more space efficient than a simple RAID1.
In the case of a RAID1, noone uses 100 mirrored drives. You use RAID10, and in the case of a failed disk, you must read 10TB to recover. With the same URE, we'd see on average 8 URE for every 10 rebuilds, or around 2 orders of magnitude less failure rate compared to the RAID6 example.
During a RAID6 rebuild, a URE is non-critical as the Array can recover the data with one lost disk an a URE on any other disk during the stripe rebuild.
The only critical error would be a URE on two disk on the same stripe, 80 URE's during a 990TB rebuild have an amazingly low chance of having two UREs on the same stripe on two seperate disks.
In case of the RAID10, you get 8 URE's over 10 rebuilds, which aren't recoverable unless you have 3 disks. So you'll corrupt data.
edit: URE of 10^14 is what most vendors specify for consumer harddrives, 10^16 is closer to what people encounter in the real world but 10^14 is considered the worst case URE rate.
URE does not have to corrupt data, if you use a proper filesystem with checksumming such as the ZFS.
When a disk fails, a RAID10 is simply in a far better position as it only have to read a single disk, and it doesn't have any complicated striping to worry about. Just clone a disk.
No but afaik there is no way to recover data once ZFS has declared it corrupted. (ie, no parity)
>The strain of rebuild have been known to kill many arrays, both RAID5 and 6.
I haven't actually encountered that yet. Despite that, a RAID 6 can loose a disk, so as long as you don't encounter further URE's after loosing another disk, it's fine.
If you're worried about that, go for RAIDZ3 or equivalent. With something like SnapRAID you can even have a RAIDZ6, loosing 6 disks without loosing data. The chances of that happening are relatively low.
>When a disk fails, a RAID10 is simply in a far better position as it only have to read a single disk
A RAID 10 is in no position to recover from URE's once a disk has failed unless you reduce your space efficiency to 33%.
I personally favor not corrupting data over rebuild speeds.
Striping might be complicated but that doesn't make it worse.
It might be acceptable too loose a music file, but once the family image collection gets corrupted or even lost on ZFS because a disk in a RAID 1 encountered a URE, it's personal.
I'd rather life with the thought that even if a disk has a URE, the others can cover for it. Even during a rebuild.
That, of course, is really not a concern for enterprise use cases.
I'd love to run ZFS, but I can't because I need to be able to add more drives as I buy them.
But it doesn't look like there's been any movement on an implementation and it seems like it's high effort and mostly home users who want this, not enterprises who might be willing to pay for.
Ah well, guess I'll stick with mdraid for now.
You could also add another raidz vdev if you have the space in the server, eventually they should level out depending on workload.
"Best" solution, read most expensive, would be to copy the data over to a new server you've configured for the new load/capacity.
Both have high investment costs for something a plain mdRAID, snapRAID or LVM RAID can achieve far simpler with better results.
A 4TB Seagate IronWolf drive costs $129 off Amazon, buying two to add as a new mirrored vdev to my TrueNAS box isn't outrageous.
I actually went custom because those off-the-shelf boxes are either very expensive or have weak CPUs so can't be used for video transcoding very well - it's cheaper to build it yourself, much cheaper if you already have old hardware to dedicate to the task.
Having an L2ARC helps out quite a bit, but only having 32GB of memory and wanting to keep most of it for the L1ARC means I still hit my spinning disks regularly (and mirrored vdev's help read IOPS tremendously in this case).
My own personal budget is very limited, buying two IronWolf HDDs in this case is easily a good chunk of my monthly income. If I can build a NAS that can expand with single drives as needed, it's more cost effective for me.
And I imagine a lot of others have the same problem.
In the end, by buying 2 drives when a single drive could have solved the problem equally well: expanding your space by 4TB, is wasting money. Period.
- Using parity rather than mirroring. I'm happy to deal with some loss of IOPS in exchange for extra usable storage.
- That deals with bitrot.
- That I can migrate to without somehow moving all of my files somewhere first (i.e. supports addition/removal of disks).
- Is stable (doesn't frequently crash or lose data)
- Is free or has transparent pricing (not "Contact Sales").
- Ideally, supports arbitrary stripe width (i.e. 2 blocks data + 1 block parity on a 6 disk array)
Unfortunately it doesn't appear that a solution for this exists:
- ZFS doesn't support addition of disks unless you're happy to put a RAID0 on top of your RAID5/6 and it doesn't support removal of disks at all when parity is involved. It is possible to migrate by putting giant sparse files on the existing storage, filling the filesystem, removing a sparse file, removing a disk from the original FS and "replacing" the sparse file with the actual disk but this is somewhat risky.
- BTRFS has critical bugs and has been unstable even with my RAID1 filesystem.
- Ceph mostly works but I always seem to run into bugs that nobody else sees.
- I couldn't even figure out how to get GlusterFS to create a volume.
- MDADM/hardware RAID don't deal with bitrot.
- Minio has hard coded N/2 data N/2 parity erasure coding, which destroys IOPS and drastically reduces capacity in exchange for an obscene level of resiliency I don't need.
- FlexRAID either isn't realtime or doesn't deal with bitrot depending which version you choose.
- Windows storage spaces are slow as a dog (4 disks = 25MB/s write).
- QuoByte, the successor to XtreemFS has erasure coding but has "Contact Us" pricing and trial.
- Openstack Swift is complex as hell.
- BcacheFS seems extremely promising but it's still in development and EC isn't available yet.
I'm currently down to fixing bugs in Ceph, modifying Minio, evaluating Tahoe-LAFS and EMC ScaleIO or building my own solution.
I've done something similar for the purpose of getting FDE with ZFS in linux. It can be a little finicky, but it's definitely workable.
One ZFS-specific caveat (which may conflict with your desire to get high storage efficiency): you way need to prevent your ZFS pools from filling up too much [1]. You can either enable discard/TRIM on the whole stack, so the top level FS (e.g. ext4) can let ZFS know when a block is actually free. Or alternative just to limit your zvols to 85% (for example) of their respective pools. The latter is my preference, because there was originally a bug with discard in zfs and it's not immediately clear if it's totally fixed (although my fstrim tests seemed to work out fine).
[1] https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_ha...
I ran some quick benchmarks (data below). Obviously this is far from rigorous, but maybe it'll be useful. In previous tests I found that volblocksize=128K was optimal for my stack -- which is why the last benchmarks use that setting.
Every additional ZFS filesystem in the stack may reduce storage efficiency (minimum free space requirements [1]; metadata & checksum overhead [2][3]) -- that's why I used ext4 as the top layer instead of another ZFS.
[1] (as mentioned before) https://www.reddit.com/r/zfs/comments/3vtur4/what_exactly_ha...
[2] https://news.ycombinator.com/item?id=14756360
[3] https://forums.freenas.org/index.php?threads/what-is-the-exa...
Test setup:
debian stable
kernel 4.9.0-3-amd64
zfs 0.6.5.9-5
ZFS "pool": mirror with 2x 7200rpm drives
Benchmark command:
for i in `seq 1 10`; do sync; dd if=/dev/zero of=DEST bs=1M count=1024 conv=fdatasync; done
zfs mirror -> dataset
Data (MB/s): 125,115,104,135,148,170,135,151,118,119
Mean (MB/s): 132.0
Std.dev.: 19.9
zfs mirror -> zvol (volblocksize=8K [default])
Data (MB/s): 150,115,127,125,122,118,105,118,124,128
Mean (MB/s): 123.2
Std.dev.: 11.6
zfs mirror -> zvol (volblocksize=128K)
Data (MB/s): 68.5,112,115,114,94.3,85.1,83.1,98.4,120,108
Mean (MB/s): 99.8
Std.dev.: 16.9
zfs mirror -> zvol (volblocksize=128K) -> luks -> ext4 (my stack)
Data (MB/s): 130,94.4,109,139,138,125,94.9,124,134,133
Mean (MB/s): 122.1
Std.dev.: 16.8
edit: formattingIn my case, I wanted a reliable filesystem, RAID1/mirror support, block-level checksumming, and full-disk encryption. No filesystem provides these on linux right now. My solution was therefore to use ZFS to provide a mirrored, checksummed, reliable volume -- onto which I put a standard LUKS-encrypted ext4 filesystem. In the past I had tried the opposite (LUKS on the bare drives, then a ZFS mirror of the 2 decrypted volumes), but it was kinda annoying to manage, and I don't really need the other ZFS features (like snapshotting).
In the grandparent post, the requirement was the ability to expand the RAID volume without rebuilding (something that ZFS doesn't offer), plus checksumming and reliability (which ZFS does offer). So one option would be to use mdadm to manage the RAID array, and then put ZFS on the resultant volume in order to get checksumming.
The disadvantages are: extra complexity; extra overhead; more potential points of failure; more management hassle; etc. As soon as encryption for ZFSonlinux is stable, I'll be very happy to drop this filesystem stacking in favor of that!
Hopefully one of the other solutions will work, otherwise I'll just have to build it.
They have a patreon if you wish to contribute that way.
At this rate XFS will end up evolving to add all the features btrfs promised before the latter makes them stable.
It may be "RAID0 on top of your RAID5/6" but it's all integrated and works smoothly. What I do is have two raidz2 vdevs of 4 disks each, one twice the size of the other, and alternate which one I upgrade (i.e. I started with 4x250gb disks and 4x500gb disks, a few years later replaced the 250gb ones with 1tb ones, then the 500gb ones with 2tb ones, and most recently the 1tb ones with 4tb ones). But yeah I am now stuck with at least 8 disks and if I wanted to migrate off them I'd have to do so all in one go.
I know that it does but I don't like the way you lose your entire pool if a single vdev fails. I'd have less of a problem with it if a pool could distribute data such that a vdev failure results in the loss of only what was on that vdev.
[1] http://www.spacex.com/about/capabilities in case you need one
It works thusly: ZFS creates a vdev inside the new larger vdev, then moves all the data from the old vdev to the new vdev, then when all these moves are done the nested vdevs are enlarged.
What should originally have happened is this: ZFS should have been closer to a pure CAS FS. I.e., physical block addresses should never have been part of the ZFS Merkle hash tree, thus allowing physical addresses to change without having to rewrite every block from the root down.
Now, the question then becomes "how do you get the physical address of a block given just its hash?". And the answer is simple: you store the physical addresses near the logical (CAS) block pointers, and you scribble over those if you move a block. To move a block you'd first write a new copy at the new location, then overwrite the previous "cached" address. This would require some machinery to recover from failures to overwrite cached addresses: a table of in-progress moves, and even a forwarding entry format to write into the moved block's old location. A forwarding entry format would have a checksum, naturally, and would link back into the in-progress-move / move-history table.
During a move (e.g., after a crash during a move) one can recover in several ways: you can go use the in-progress-moves table as journal to replay, or you can simply deref block addresses as usual and on checksum mismatch check if you read a forwarding entry or else check the in-progress-moves table.
For example, an indirect block should be not an array of zfs_blkptr_t but two arrays, one of logical block pointers (just a checksum and misc metadata), and one of physical locations corresponding to blocks referenced by the first array entries. When computing the checksum of an indirect block, only the array of logical block pointers would be checksummed, thus the Merkle hash tree would never bind physical addresses. The same would apply to znodes, since they contain some block pointers, which would then have three parts: non-blockpointer metadata, an array of logical block pointers, and an array of physical block pointers.
The main issue with such a design now is that it's much too hard to retrofit it into ZFS. It would have to be a new filesystem.
Huh?
> IDK if that's in OpenZFS yet.
The openzfs tree (on github) is virtually identical to illumos-gate (on github).
> physical block addresses should never have been part of the ZFS Merkle hash tree, thus allowing physical addresses to change without having to rewrite every block from the root down.
mahrens deals with this (and block pointer rewriting) here:
https://www.youtube.com/watch?v=G2vIdPmsnTI#t=44m53s
Even with SSDs IOPS are precious. On rotating media, burning track-to-track seeks in reading and updating a large hash table is a bad plan (cf. the deduplication table).
It was not a per se data loss bug. It was Btrfs corrupting parity during scrub when encountering already (non-Btrfs) corrupted data. So a data strip is corrupt somehow, a scrub is started, Btrfs detects the corrupt data and fixes it through reconstruction with good parity, but then sometimes computes a new wrong parity strip and writes it to disk. It's a bad bug, but you're still definitely better off than you were with corrupt data. Also, this bug is fixed in kernel 4.12.
https://lkml.org/lkml/2017/5/9/510
Update, minor quibbles:
lacking in Btrfs is support for flash Btrfs has such support and optimizations for flash, the gotcha though if you keep up with Btrfs development is there have been changes in FTL behavior and it's an open question whether or not these optimizations are effective for today's flash including NVMe. As for hybrid storage, that's the realm of bcache and dm-cache (managed by LVM) which should work with Btrfs as any other Linux file system.
ReFS uses B+ trees (similar to Btrfs) XFS uses B+ trees, Btrfs uses B-trees.
According to some bug reports, nobody has touched this since 2011...
I recently set up a ZFS volume using 12x4TB drives using RAID-Z2, so I expected 40TB of usable space, or ~36.3TiB. However, I only see 32TiB of usable space on the volume. I always wondered why that was so, never figured it out..
Basically, ashift=12 increases ZFS block size to 4K. Metadata use full blocks that would be 512b on ashift=9 but are now 4K (due to ashift=12). It wastes at least 3.5Kb more than normal 512 byte blocks for each block that is not filled entirely.
Incidentally, many modern filesystems (including NTFS) store very small files in the FAT rather than taking up a whole block for this very reason!
ZFS has this same feature, however, as long as feature@embedded_data=enabled ;)
Best is to just call it metadata :)
Trying to do this with FS features is misguided.
You need to have backups, and have regular practice in restoring from backups.
Some organizations need fancy filesystems in addition to backups, because they want to have high availability that will bridge storage failures. But that has a high cost in complexity, you should only consider it if you have IT/sysadmin staff and the risk management says it's worth the investment in cognitive opportunity cost, IT infrastructure complexity and time spent.
Yes, there is still a risk that corrupted data may end up in backups, but that's true even with ZFS. Ideally you want end-to-end integrity checking and verification, that means application layer and should also be done for backups. But like with all risk management, there are diminishing returns...
After years of btrfs I realized while the all the features around snapshotting, send/receive etc are great the cost in performance and other issues is too high.
And using plain old ext4 is more often than not the best compromise so you can forgot just about the fs and focus on higher layers.
Snapshots are nice, sure. But i'd rather do that on top of my filesystem (they're called backups) and leave the filesystem lean, simple, fast AND reliable. Every filesystem may have its own problems, but i feel that the "attack surface" of my filesystem should be as small as possible. Do i rather want to have a bug in my filesystem implementation concerning snapshots or would i rather have that bug in my backup tool?
IMO, the FS should be rockstable and lean and not "cool and fancy". I can do fancy on top of rockstable.
If your use case mainly revolves around the benefits of snapshots then it definitely makes sense.
It seems that the GP's point that the fs should just store files is rather debunked, as snapshots are an extremely useful feature for a filesystem.
https://access.redhat.com/documentation/en-US/Red_Hat_Enterp...
https://clearlinux.org/blogs/linux-os-data-compression-optio...
Though it looks like zstd is coming, which is exciting: https://reviews.freebsd.org/D11124
Or use lz4 while you wait.
I use nas4free with much less ram…
Between the low cost of storage, and alternative solutions for deduplicating data I personally don't use the built-in deduplication functionality of ZFS for my zpools. Might come down to what sorts of data you are storing, though.
Over time it would become horrendously slow, I agree, since you have too little RAM by a factor of at least 8.
I've got 30TB of small files stored on our TrueNAS system at work, there's no benefit to enabling it but if someone decided to toggle it we'd quickly learn there isn't enough RAM in the world to handle billions of 1-16KB files....
Want to test the upgrade in your production environment beforehand? Well, make a clone a couple days early.
And once all works out, a few days later you promote the clone and destroy the old datasets. Need a rollback? Well, just start the old application instance on the old datasets, nothing touched them.
Doesn't work for all database types, especially if you have no possibility to replay new data into the rollback. But if your system allows it, it is really comfortable.
As long as you make sure to only use one filesystem, i.e. you don't place pg_xlog or some tablespaces on a different filesystem. You can get very weird corruption in such cases :)
Bitrot occurs because the lower-abstraction level hardware fails. When you put a Hard-Drive into storage for say 5 years, the bits may change. Even if a Hard-drive remains in constant use for 5 years... if said files or directories aren't checked and double-checked constantly, the error-correction codes may fail over time.
Its a fundamentally different problem from Hard Drives that are being used constantly as say Swap.
Hard Drives typically include Hamming codes or ECC bits to address typical corruption issues.
-------------
The fundamental principle at hand here is as follows: to ensure integrity of files, you need to regularly check file data. Only the Filesystem would know which files were recently checked.
The "Bit-Rot" scenario is particularly harmful to RAID5 (Minimum 3-hard drives. Two contain data, one contains "parity" that can fix any errors on the other hard drives. Then the parity is structured to be striped equally across the three drives). Modern RAID drivers can do this rather easily.
The problem with "Bit Rot" is that a RAID5 drive will not rebuild itself until it detects an error. If you're reading files along, and all of a sudden... the hard drive detects an error. No problem (in the typical case), just rebuild the data from the parity.
However, "Bit Rot" means that the parity bits (on the 3rd backup hard drive) have ALSO rotted away.
----------
The only way to fix this "bit-rot" error is to constantly read through your data and CONSTANTLY check for bit-rot. No hard drive is going to silently spin and hamper-performance of the system for self-verification purposes... but a Filesystem / Operating system can schedule these "Scrubs" to occur during periods of low-I/O.
Which is how ZFS, and Window's ReFS work. When your computer is idle, the OS checks for bitrot. When the computer starts to work again, it pauses the "low priority" bitrot checks and serves the data.
-------
ZFS doesn't quite work like Window's ReFS. ZFS simply checks for bit rot whenever a file is accessed. Every time. There are "ZFS Scrub" commands (which you can put into a cron-job) to read every file (and therefore check for bitrot).
You have a glob of storage that is presented as a block device, but actually underneath is a reas no fooling filesystem with FEC, Snapshots and all sorts of other goodies.
It's been posited that the main push for ZFS was perceived bugs in UFS that happened to be LSI firmware bugs once they had the capabilities of ZFS to detect them.
I mean, as it is now it seems like we have a hard enough time dealing with comparatively simple hybrid memory systems.
http://files.gpfsug.org/presentations/2016/anl-june/LANL_GPF...
> And every enterprise system has already moved way past what ZFS can do, including enterprise-class offerings based on ZFS from Sun, Nexenta, and iXsystems.
However, OP is naturally suggesting ZFS over NTFS, HFS+, ext3/4 and even ReFS and APFS.
> Still, ZFS is way better than legacy storage SOHO filesystems. The lack of integrity checking, redundancy, and error recovery makes NTFS (Windows), HFS+ (macOS), and ext3/4 (Linux) wholly inappropriate for use as a long-term storage platform. And even ReFS and APFS, lacking data integrity checking, aren’t appropriate where data loss cannot be tolerated.
Basically every storage array is years beyond ZFS in terms of features and capabilities. Even the ones based on ZFS! And I don't think anyone would argue that. After all, Dell EMC, HPE, Pure Storage, HDS, etc spend many millions a year developing advanced storage technologies.
ZFS is a filesystem (and also kind of a volume manager) and is designed to be fairly versatile and to run alongside an OS and applications. ZFS includes some impressive features long found in enterprise arrays, including snapshots, mirroring, replication, RAID, and data protection.
Dedicated storage arrays have long supported many more specialized features, and I urge you to read up on cool products like Pure FlashArray, NetApp SolidFire, EMC Unity, EMC Isilon, HPE Nimble Storage, Tintri, etc. You'll see that most/all have awesome hybrid (flash/disk) capabilities, all-flash tuning, integration with VMware, Windows Server, and Docker, scale-out capability, etc.
It's a pretty amazing field, and that's why I've dedicated my professional career specializing in enterprise storage. Maybe a good place to start is the Storage Field Day video series on YouTube: https://www.youtube.com/results?search_query=%22Storage%20Fi...
This is why NetApp sued Sun in 2007, claiming patent infringement.
I worked in a data center that stored medical records around that time. We were moving off of a StorageTek PowderHorn 9310 (that thing was mesmerizing) to some multi-cabinet EMC spinning disk array. You couldn't get access to that thing without hooking up a FC HBA to the SAN. Fast-forward a few years and I found myself in a similar job at a much smaller e-commerce firm. We had some old finicky NetApp filers we ended up getting rid of in favor of Sun Storage 7000 storage appliances running ZFS. Both of them supported replication, but the Sun Storage appliances were far better.
Yeah, SAN and NAS are two different things, but there's a lot of money, a lot of shared technology, and a lot of similar use cases that have driven convergence over the last decade or two.
As for the rest, I was looking for more concrete comparisons, and third-party software integrations don't seem relevant to this conversation.
basically lots of nice python wrapper around GPFS.
ZFS isn't clustered and has no concept of HSM, so migrating hot blocks from NVMe to SSD then disk and tape isn't an option
Also stuff around disk rebuilding is still pretty old. Modern FSs can detect a failed disk and rebuild the array in minutes based on FEC and even distribution of data.
The ZFS storage appliance offers clustered configurations: https://www.oracle.com/storage/nas/index.html
It's unclear what your requirements are when stating "clustered".
and has no concept of HSM, so migrating hot blocks from NVMe to SSD then disk and tape isn't an option
Oracle HSM integrates with ZFS:
https://www.oracle.com/storage/tape-storage/hierarchical-sto...
Also stuff around disk rebuilding is still pretty old. Modern FSs can detect a failed disk and rebuild the array in minutes based on FEC and even distribution of data.
ZFS can detect a failed disk and rebuild in minutes as well depending on storage and system speed, etc. so I'm wondering what you're referring to?
ceph is horrifically slow, and only has a rudimentary posix interface.
gluster is just well, terrible.
first things first a SOHO office cannot support a clustered filesystem, unless one of the people happens to be a storage specialist.
yea I hear lots of noise about how self healing they both are. That is mostly fancy talk for "I don't have backups"
Supposedly they both do HA. But yeah, its not something I'd want to support.
If you look at gitlab's setup, and their proposed setup (https://about.gitlab.com/2016/12/11/proposed-server-purchase...)
32 file servers to serve ~480TBs of disk. Seriously? 4 file servers, GPFS and 4 MD3060e. Thats about a petabyte of usable storage with a streaming throughput of about 40 gigabits a second.
"modern" clustered filesystems are mainly just toys. if you want speed, use lustre, if you want sexy software defines awesomeness with next level raid, use GPFS.