Demystifying OpenZFS 2.0
klarasystems.com
klarasystems.com
My background includes a few years at IBM as an OS architect in charge of, among other things, the file systems and network services for AIX. However, this was a long time ago (~1988). So I understood NFS, inodes, what it meant to mount a device and the classic Unix commands for managing filesystems.
The first ZFS problem for me was just the vocabulary. What's a pool? What's a VDev? Fortunately, there are lots of high-level descriptions of ZFS on the internet that cover the high level concepts.
Getting the hardware set up was the second hurdle. There are a lot of discussions on forums that cover hardware issues and there you will find answers to many hardware issues, like: is ECC memory required, how fast does the processor need to be, are SATA drives recommended and what drive controllers are compatible with ZFS. Naturally, the answers to these questions depend on how critical is the system you are putting together. I suggest for most people to try setting up a ZFS on some old hardware just to get practice with it. I had a low end Dell tower server model that I wasn't using so it became my first FreeBSD/ZFS server. It had four hard drives.
The real difficulty for me was just the lack of a deep understanding of how ZFS works. ZFS has lots of features totally new to me so working with only the man pages for the ZFS related commands was not giving me much fun. I ended up buying two small books that I can't recommend enough, FreeBSD Mastery: ZFS and FreeBSD Mastery: Advanced ZFS by Michael W Lucas and Allan Jude, see [1] and [2].
These books are each around 200 pages long. The chapters of the first book are:
1. Introducing ZFS
2. Virtual Devices
3. Pools
4. ZFS Datasets
5. Repairs and Renovations
6. Disk Space Management
7. Snapshots and Clones
8. Installing ZFS
Each of the chapters is quite detailed, and I found these two books were just what I needed.
[1] https://read.amazon.com/kp/embed?asin=B00Y32OHNM&preview=new...
[2] https://read.amazon.com/kp/embed?asin=B01E40YIRM&preview=new...
I applaud OpenZFS for these major improvements, but its performance is not competitive with some of the commercial offerings, even if OpenZFS is a more easily administered solution.
The main issue with OpenZFS performance is its write speed.
While OpenZFS has excellent read caching via ARC and L2ARC, it doesn't enable NVMe write caching nor does it allow for automatic tiered storage pools (which can have NVMe paired with HDDs.)
This means that large writes, those which exceed the available RAM, will drop down to the HDD write speed quite quickly instead of achieving wire speed. Which means I get write speeds of 300 MBps versus 10 GBps wirespeed.
The solutions competitors have implemented are:
- SSD read-write caching, usually using mirrored NVMe drives to ensure redundancy.
- SSD tiering, which is including your SSDs in your storage pool. Frequently accessed data is moved onto the SSDs while infrequently accessed data is moved to the HDDs.
Without these features, my QNAP (configured with SSD tiering) kicks the pants off of ZFS on writes that do not fit into memory.
References:
[1] QNAP Tiering: https://www.qnap.com/solution/qtier/en-us/
[2] QNAP SSD Read-Write Caching: https://www.qnap.com/solution/ssd-cache/en/
[3] Synology SSD Read-Write Caching: https://kb.synology.com/en-my/DSM/help/DSM/StorageManager/ge...
I'm assuming that QNAP are using ZFS, which a cursory search supports.
QNAP has two NAS OSes right now. The newer one is QuTS and it support ZFS. The other older one is QTS which doesn't use ZFS. Only the older QTS OS supports SSD Tiering.
> Is QNAP doing something outside the ZFS spec here? If not, it's almost certain that solutions like TrueNAS can replicate this behaviour. If they are, I'd be concerned about vendor lock-in and data recovery.
I am using TrueNAS for my ZFS solution and it doesn't offer anything out of the box to address the issues I am bringing up.
I would suspect that QNAP is using one of them.
You are spreading incorrect information, please stop. You know enough to be dangerous, but not enough to actually be right.
>While OpenZFS has excellent read caching via ARC and L2ARC, it doesn't enable NVMe write caching nor does it allow for automatic tiered storage pools (which can have NVMe paired with HDDs.)
Huh? What are you talking about? ZFS has had a write cache from day 1: ZIL. ZFS Intent Log, with SLOG which is a dedicated device. Back in the day we'd use RAM based devices, now you can use optane (or any other fast device of your choosing including just a regular old SSD).
https://openzfs.org/w/images/c/c8/10-ZIL_performance.pdf
https://www.servethehome.com/exploring-best-zfs-zil-slog-ssd...
The ZIL is specifically not a cache. It can have cache-like behavior in certain cases, particularly when paired with fast SLOG devices, but that's incidental and not its purpose. Its purpose is to ensure integrity and consistency.
Specifically, the writes to disk are served from RAM[1], the ZIL is only read from in case of recovery from an unclean shutdown. ZFS won't store more writes in the SLOG than what it can store in RAM, unlike a write-back cache device (which ZFS does not support yet).
[1]: https://klarasystems.com/articles/what-makes-a-good-time-to-...
So many people misunderstand the ZIL SLOG, including the guy you are responding to.
I think most posts on the internet about the ZIL SLOG do not explain it correctly and thus a large number of people misunderstand it.
There was some work done[1] on proper write-back cache for ZFS by Nexenta several years ago, but it seems it either stalled or there was a decision to keep it in-house as it hasn't made its way back to OpenZFS.
Exactly, and it's even worse. By default ZFS only stores about 5 seconds worth of writes in the ZIL, so if you have a NAS with say a 10Gbe link that's less than 10GB in the ZIL (and hence SLOG) at any time.
> A real write cache would be awesome for ZFS.
Agreed. It's a real shame the write-back cache code never got merged, I think it's one of the major weaknesses of ZFS today.
So if you really wanted to, you could try looting code from there, though I believe Nexenta's implementation of that was either orthogonal or significantly different from the one that landed in OpenZFS...
[1] - https://github.com/Nexenta/illumos-nexenta/blob/release-5.3/...
[2] - https://github.com/Nexenta/illumos-nexenta/blob/release-5.3/...
You are spreading incorrect information, please stop. You know enough to be dangerous, but not enough to actually be right.
I've benchmarked ZIL SLOG in all configurations. It doesn't speed up writes generally. It speeds up acknowledgements on sync writes only. But it doesn't act as a write through cache in my readings and in my testing.
What it does is it allows a sync acknowledgements to be sent once a sync write is written to the ZIL SLOG device.
But it doesn't actually use the ZIL SLOG for reading at all in normal operation, instead it uses the in-memory RAM to cache the actual write to the HDD-based Pool. Thus you are still limited to RAM size when you are doing large writes -- you may get acknowledges quicker on small sync writes, but your RAM limits the size of large write speeds because it will fill up and have to wait for it to be saved to the HDD to accept new data.
Here is more data on this:
https://www.reddit.com/r/zfs/comments/bjmfnv/comment/em9lh1i...
e: Not doing market research here, I'm just seeing (sometimes, feeling) all the pain involved with cluster filesystems and wondering how your solution stacks up.
Yes, nodes have NVMe SSD local scratch space for fast compute needs. Large datasets reside in central storage. Smaller chunks are usually brought to local scratch to let compute jobs do their thing. Results go back to central storage again.
I built it for our needs based on freely available docs and community support. The server runs FreeBSD + ZFS + NFS and nothing else. Snapshots from this primary storage gets shipped over to a secondary location that has a slow mirror(similar setup but with cheap SATA disks). Our need is only minor geographical data durability and not lower downtime DR.
Experience: beats most commercial stuff we have run so far. Minimal setup of FreeBSD, ZFS, and NFS runs buttery smooth with almost no operational needs. I run upgrades on OS after it has been in community for a while, and it has only been smooth. We didn’t find a need to buy specialist support as our setup is minimal, and workload not any crazy, and availability targets not so strict.
Performance wise this setup far exceeded our expectations in all our usecases. We do have 768 GiB of memory for this that ZFS happily makes use of. ARC is about 60-70% of mem most of the time.
I will happily build the same again should there be a need for storage.
So, we make do with what we have. We have three similar storage configs(and three mirrors as slow backups). We carry spare memory sticks, Optane drives, disks, and power supply for minimising down time due to hardware failures. If something were to go wrong for which we don't have spares, yes, the storage will be offline until we can fix it.
Planned updates? - we have two scheduled downtimes in a year. All patches to these machines happen at that time. These machines are strictly internal and don't have internet access except for those planned maintenance days.
No, the ZIL is just for integrity[1].
Sure if your workload is primarily long, sustained transactions the SLOG will do little. But for any workload that is even a little bursty in nature (like hosting virtual machines) it will absolutely make a difference.
All of the above is assuming you have enabled sync=always, which I hope people are doing for production workloads unless they happen to have an app that can tolerate lost writes/partial writes.
As such it can behave a bit like a write-back cache as I mentioned in the linked post.
A crucial difference is that the ZIL is limited by the available memory regardless (and tunables, by default 5 seconds worth of data), which would not be the case for a write-back cache device.
1. I didn't ever call it a write-back cache and never even hinted at it operating like one
2. Write-back caching is only one of many options for write caches
Regardless of both of those points, the SLOG ABSOLUTELY will speed up writes if you have a bursty workload (like the virtual machines he's running) and sync=always enabled.
>A crucial difference is that the ZIL is limited by the available memory regardless (and tunables, by default 5 seconds worth of data), which would not be the case for a write-back cache device.
And yet that point is irrelevant to OPs question of whether or not SLOG was speeding up his writes, which it almost assuredly is doing with the amount of memory and the size of his SLOG device when he is doing sync writes.
The point remains though. The ZIL is not meant to be a cache. The SLOG (which is a "ZIL on a separate device") could be viewed as a cache, but it's a cache for the ZIL. Thus being a cache for writes is incidental, not the primary purpose.
Certainly, if your workload aligns with those conditions, then you will experience a performance boost. For other workloads it it won't help and for certain workloads it can even be detrimental[1].
[1]: https://www.reddit.com/r/zfs/comments/m6ydfv/performance_eff...
ZIL and SLOG are NOT a write cache. They are 100% about not loosing data. A ZIL is part of a zpool where data is written to when the file protocol requires sync. When it is written the sync is cleared. In this case you are sharing the zpool with all its others needs to write and read, so more load on the pool. The data is also in RAM given to the write aggregator process. The RAM is only written latter. If the system crashes before the aggregator writes the data, on restart the data written to the ZIL is read into the aggregator and then is written out to the zpool. By default the ZIL is never read. An SLOG is just a ZIL that is on a separate device from the zpool. So a sync write comes in and is written to the SLOG and not part of the spinning rust zpool so no shared resources with the pool. The sync is marked as cleared when the data is written to the SLOG. The SLOG should be really fast as the goal is to ack the sync quickly. Remember at this point the data is NOT written to the zpool, it’s RAM and the aggregator will write the data out later. Same thing applies on crash. Reboot, SLOG is reviewed for uncleared items, sent to RAM, written to pool.
TLDR - zil/slog are write devices where the data is written when sync is required and almost never read back. Data is alway in RAM cache and written out to the spool from RAM. ZIL/SLOG is not a buffer. Speed of ack sync will be as fast as data can be written to ZIL/SLOG.
Case A : No SLOG. You have sync writes and they hit memory, the transaction groups get flushed to disks and _only then_ the ack goes back completing the write IO operation.
Case B : low latency SLOG. You have sync writes and they hit memory, immediately gets written to SLOG, immediately the ack goes back completing the write IO operation while batching in transaction group does the flush to disks in background(at slow disk speeds). Sure, if you are out of memory to commit those write IOs, you're going to hit the disks in hot path and get undesirable perf.
Case B surely is always better than case A wouldn't you say?
In my scenarios were I ran out of memory on large writes and them get slow downs I would love a real write cache on an NVMe or tiered storage. It would make a huge difference.
At first I thought you were talking about actual raw read/write speed and how things like ARC or write caches can actually become bottlenecks when using NVMe storage, which can easily get to 200 Gbps with a small number of devices. That's being worked on via efforts like adding Direct IO.
Instead though I think you've fallen into one of those niche holes that always takes longer to get filled because there isn't much demand. Big ZFS users simply have tons and tons of drives, and with large enough arrays to spread across even rust can do alright. They'll also have more custom direction of hot/cold to different pools entirely. Smaller scale users, depending on size vs budget, may just be using pure solid state at this point, or multiple pools with more manual management. Basic 2TB SATA SSDs are down to <$170 and 2TB NVMe drives are <$190. 10 Gbps isn't much to hit, and there are lots of ways to script things up for management. RAM for buffer is pretty cheap now too. Using a mirrored NVMe-based ZIL and bypassing cache might also get many people where they want to be on sync writes.
I can see why it'd be a nice bit of polish to have hybrid pool management built-in, but I can also see how there wouldn't be a ton of demand to implement it given the various incentives in play. Might be worth seeing if there is any open work on such a thing or at least feature request.
Also to your lower comment:
>I am using TrueNAS for my ZFS solution and it doesn't offer anything out of the box to address the issues I am bringing up.
While not directly related to precisely what you're asking for, I recently started using TrueNAS myself for the first time after over a decade of using ZFS full time and one thing that immediately surprised me is how unexpectedly limited the GUI is. TONS and tons of normal ZFS functionality is not exposed there for no reason I can understand. However, it's all still supported, it's still FreeBSD and normal ZFS underneath. You can still pull up a shell and manipulate things via the commandline, or (better for some stuff) modify the GUI framework helpers to customize things like pool creation options. The power is still there at least, though other issues like inexplicably shitty user support in the GUI show it's clearly aimed at home/SoHo, maybe with some SMB roles.
You are correct that prices of small SSD are near HDD prices, but most people would want larger drives like 12 to 16TB drives in their array, and SSDs can not compete with these on price at all.
Maybe we are in the twilight of HDDs? And SSDs will compete on price across the whole storage capacity range soon? Maybe...
if you needed to write a LOT (let's say something like a sustained 70MB per second all the time) but the data rate was not particularly high, you might easily wear out a RAIDZ2 composed of 2TB-4TB sized cheap SSDs and kill them in a fairly short period of time, when the same RAIDZ2 composed of 2.5" 15mm height 5TB spinning drives, or 3.5" 14-16TB drives would successfully run for many years.
also re: cheap SSDs, the quad level cell tech and write cache is a big limitation in sustained write speeds:
https://www.firstpost.com/tech/news-analysis/samsung-870-qvo...
That is incorrect.
RaidZ(x) vdev write speed is equal to the write speed of a single drive in the vdev.
If you want a raidZ zpool to have faster writes, you add a second vdev - that would double your write speed.
A wider vdev, on the other hand (15 or 18 drives, etc.) would not have faster write speeds.
I may be heavily influenced by my own situation, but in France, SSDs still are much more expensive than HDDs. I've just bought four IrownWolves, €100 for 4 TB. The cheapest 2 TB SSD I can remember was around €180.
My point is that given the price and the usage profile, I kind of agree with the other poster: I don't have a use for tiered ZFS and would much rather they spent the time on other things.
As it is, I'm already doing a tiered storage of sorts: data that I'm currently working with requiring fast i/o is directly attached to my computer. "Cold" or less i/o sensitive data is plenty fast on spinning drives across the network.
As a home user living in an apartment, this affords another optimization: kick the NAS out and access it over the internet, because gigabit speeds become sufficient.
You sure wouldn't want to use them in any write heavy zfs application however, because you'll kill them quick in their ultimate TB write endurance. Those are triple level and quad level cell SSDs. Double the dollar figures there for anything with a ultimate lifespan TB written figure that you would want to try in a ZFS array.
> Synology
I am perfectly willing to sacrifice some performance in exchange for not being locked into a vendor-specific proprietary NAS/SAN hardware platform or closed source software solution that's proprietary to that one vendor.
That is to say that yes, it does have a bunch of proprietary stuff on top, but it also:
1. Runs Linux
2. Uses Linux systems, like BTRFS and LVM, to pool data
3. Provides SSH access
4. Has tons of packages that you might want to use (Squid, Node, Plex, etc.) available
5. Supports Docker, among other things, so you can run what you like.
So yeah, they have a lot of proprietary offerings on top of a generic Linux storage solution, but if you change your mind you can take that generic Linux storage solution somewhere else without being locked in to anything Synology provides. The only caveat is if you choose to use their all-the-way proprietary systems, like their Google Docs equivalent or their cloud music/drive/etc. stuff, all of which is still stored on the same accessible drives. You might lose your "music library" or "photo library" or "video library", but your music and photo and video files will still be accessible.
Also, since a reasonable number of people may use ZFS via TrueNAS (formerly FreeNAS), note the GUI there doesn't expose a ton of useful features itself but can still be told to do them since it handles interfacing with ZFS via a plain Python. Currently, GUI pool/fs creation stuff is handled via:
/usr/local/lib/python3.9/site-packages/middlewared/plugins/pool.py
'options' handles pool flags and 'fsoptions' fs ones. Modifying that (reboot seems needed to make it take effect afterwards though perhaps reloading something else would do it) will then allow all the GUI setup to work and slice things the way TrueNAS wants while still offering finer grained property control. Particularly useful for things that can only be set at creation time, like pool properties such as ashift/feature flags or fs properties like utf8/normalization.----
0: https://jrs-s.net/2018/08/17/zfs-tuning-cheat-sheet/
1: https://github.com/openzfs/zfs/pull/9735#issuecomment-570082...
Klara systems is more of a development and support consultancy firm for FreeBSD and ZFS.