My 2023 all-flash ZFS NAS (Network Storage) build
michael.stapelberg.ch
michael.stapelberg.ch
Edit: Looked again and he's getting redundancy by running multiple NASes and rsyncing between them. Still seems like a risky setup though.
> Corrective "zfs receive" (#9372) - A new type of zfs receive which can be used to heal corrupted data in filesystems, snapshots, and clones when a replica of the data already exists in the form of a backup send stream.
Are there any other options worth looking at? Thanks!
My homelab with Turing Pi 2 is running RPi4 CM and an Nvidia Jetson. Each also have a small SSD (one native, two miniPCI to SATA, and one USB). They could be used with k8s or a clustering filesystem but haven't played with it as of yet.
if you want speed: lustre
if you want anything else: gpfs.
Depends. At my last job we used it for our OpenStack block/object store and it was performant enough. When we started it was HDD+SSD, but after a while (when I left) the plan was to go all-NVMe (once BlueStore became a thing).
Haven't tried all-SSD cluster yet though, only spinning rust.
It pains me to see a ZFS pool with no redundancy because instead of being "blissfully" ignorant of bit rot, you'll be alerted to its presence and then have to attempt to recover the files manually via the replica on the other NAS.
I appreciate that the author recognizes what his goals are, and that pool level redundancy is not one of them, but my goals are very different.
It's going to be less optimal and less professional. That's ok, for something as important as backups, keep it boring. Simply starting up the second box is simple, stupid and has well-understood failure modes. Maybe someone like me should just buy an off-the shelf NAS.
That way everything has 3 copies, but I'm not backing up a compressed backup that might not deduplicate well and might be a little excessive to have 4 copies.
The list of open issues on GitHub in the native encryption component is quite telling: https://github.com/openzfs/zfs/issues?q=is%3Aopen+is%3Aissue...
While it’s not common, the amount of people running into edge cases and killing both their sending and sometimes even receiving pools (better have another backup!) is frankly unacceptable. “Raw sends” seem to be especially at risk, though sending in general seems to be where the issues mostly lay. My thoughts mirror the comments here: https://discourse.practicalzfs.com/t/is-native-encryption-re...
Here’s a currently unmaintained document by Rincebrain that he was using to try to track things before he got burned out by the lack or response. https://docs.google.com/spreadsheets/d/1OfRSXibZ2nIE9DGK6sww...
But I’m concerned reading these comments. Anyone else experiencing issues?
The utilities and tooling surrounding encryption are also weak, and there are ways you can throw away critical, invisible keydata without realizing it, and no tool to allow the correction of the issue, even if you have the missing keydata on another system.
It's a shame, because the main draw of the native encryption to me is being able to have a zero-trust backup target with all the advantages of zfs send over user-level backup software. But I've heard of several people running into issues doing this (though luckily no actual data loss that I've heard, just headaches)
I run the 50 zfs disks in my house on LUKS.
> Learn about how ZFS is being adapted to the ways the rules of storage are being changed by NVMe. In the past, storage was slow relative to the CPU so requests were preprocessed, sorted, and coalesced to improve performance. Modern NVMe is so low latency that we must avoid as much of this preprocessing as possible to maintain the performance this new storage paradigm has to offer.
> An overview of the work Klara has done to improve performance of multiple ZFS pools of large numbers of NVMe disks on high thread count machines.
[…]
> A walkthrough of how we improved performance from 3 GB/sec to over 7 GB/sec of writes.
* https://www.youtube.com/watch?v=v8sl8gj9UnA
When ZFS was created spinning rust was still the main thing, and SSDs were gaining popularity, so ZFS created "hybrid storage" pools:
* https://cacm.acm.org/magazines/2008/7/5377-flash-storage-mem...
* https://web.archive.org/web/20080613124922/http://blogs.sun....
* https://web.archive.org/web/20080615042818/http://blogs.sun....
* https://www.brendangregg.com/blog/2009-10-08/hybrid-storage-...
* https://en.wikipedia.org/wiki/Hybrid_array
Still useful for use cases that lean towards bulk storage (versus IOps).
Coming back on topic, the idea of having 3 custom made NASes is surely interesting, but apart from learning and experimenting with all these, I don't see a very big advantage from backup/security point of view compared to commecrial NASes (Synology, QNAP).
For sure we can all argue here about selection of filesystem (ZFS, btrfs, ext4...), selection of CPU, RAM type and all the others, but it all boils down to what each one wants (and has the money to spare). IMHO I wouldn't go with QVO SSDs and non redundant volumes (especially spanning 2 disks), but hey that's just me :-)
I see there are tri-mode backplanes but looking for something that skips the hba and connects via something like occulink directly?
I have spent an enormous amount of time over the past couple of weeks tuning ZFS to give us the best balance of reads-vs-writes, but the biggest problem is trying to find the right benchmark tool to properly reflect real-world usage. We are currently using FIO but sheer number of options (depth queue, numjobs, libaio vs io_uring) makes the tool unreliable.
For example, comparing libaio vs io_uring with the same options (numbjobs, etc) makes a HUGE different. In some cases, io_uring gives us double (or more) performance than libaio, however, io_uring can produce numbers that don't make any sense (eg: 105GB/sec reads for a system that maxes out at 72B/sec). That said, we were able to push > 70GB/secs large-block reads (1M) from 12x NVMe drives which seems to validate ZFS can perform well on these servers.
OpenZFS has come a long way from the 0.8 days, and the new O_DIRECT option coming out soon should give us even better performance for the flash arrays.
The fact that fio has a bunch of options doesn’t make the tool unreliable. Not understanding the tool or what you are testing makes you unreliable as a tester. The tool is reliable. As you learn you will become a more reliable tester with it.
I design similar NVMe-based ZFS solutions for specialized media+entertainment and biosciences workloads and have put massive time into the platform and tuning needs.
Also think about who will be consuming data. I've employed the use of an RDMA-enabled SMB stack and client tuning to help get the best I/o characteristics out of the systems.
For general video and media workloads, it may be something like, "we have to accommodate 40 editors working over 10GbE (2 x 100GbE at the server) and minimize contention while ingesting from these other sources".
I work with iozone to establish a baseline. I also have a "frametest" utility that helps when mimicking some of the video characteristics.
Not sure about current status?
https://github.com/openzfs/zfs/issues/7734
https://github.com/openzfs/zfs/issues/342
https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSForSwapM...
Swap delays that but extra RAM delays it too. If you take a use case where 2GB of memory is fine, and give it 8GB, then you already solved the problem swap would solve. You can always add more but you're past the point of diminishing returns.
But I wouldn't say it's questionable. On a NAS I doubt you'd ever use even one full gigabyte of swap space. Keeping it simple is fine.
Ed: I see he doesn't - but could (and optionally with a slice for swap partition): https://openzfs.github.io/openzfs-docs/Getting%20Started/Ubu...
BTW, currently working around a bigger one that will run on a faster mini PC and a 8 bay USB3.1 enclosure. I'm a little wary about USB connectors in this context, and I'll likely have to secure cables firmly to avoid wearing and accidental pulls, but so far results on the bench are promising.
ZFS needs some amount of ram in order to even load the pool, this only becomes a practical concern when your pool gets into the 100’s of TiB.
Deduplication, which should generally not be used, used to be awful with ram consumption, but these days isn’t nearly as bad and can be surprisingly viable. Dedupe can still cause unexpected performance issues unless your skillset is into digging into system analysis and tuning.
You need some amount of RAM buffer for the write TXGs to coalesce efficiently. Generally not a concern.
Finally there’s ARC, which is where all the nice things that improve your experience happen. The more the better, but just like system ram, once you have enough for your usage profile, you stop noticing much benefit beyond that. For dumb file storage, not that much is needed. Ideally you want enough to keep all the metadata, plus whatever your actual repeated read access would be. Working on video editing and VMs would require more RAM for a more optimal experience.
All in all I'm happy with it. I still use spinning rust for the backup though, no sense worrying about that. At the time it was cheaper to get 12x4 TB budget NVMe and the extra parts than go for any reasonable count of SATA SSDs, not sure if that's still true or not. SATA definitely would have been easier, e.g. stuff like in the Flashstor 12x2TB build I had to submit a kernel patch because the cheap drives were duplicating their nsids.
IIRC there was some talk that RAID-Z lead to write amplification compared to mirrored drives, thus not good for cheaper SSDs. Haven't had time to sit down and think it through, does anyone know if thats right or wrong?
> I used 3 4 way switches from aliexpress
Did you find any that were significantly cheaper than ~100 USD?
I want to say that sounds about right on price. If I didn't also have a desire for high single core performance at the same time a used/old Threadripper/Epyc build and bifurcation would probably make more sense. I also disconnected the onboard tiny "definitely going to be noisy as hell in 3 months" low quality fans from them and just rested a 140mm blowing down across the top of the 3 cards at 30% speed. Temps of the controllers and SSDs became better and it's dead silent.
I have a bunch of unused M.2 drives from decommissioned servers that I'd love to use as additional storage.
If you don't specifically want the high per core performance of something like a 13900k a used/old Threadripper or Epyc system and bifurcation might make more sense. It'll also enable you to get maximum per drive bandwidth, if that's a concern for your (when you have 12 drives in some form of stripe and parity the per drive bandwidth ends up not being that important for most sane workloads though).
There's also dm-integrity [1]:
> The dm-integrity target emulates a block device that has additional per-sector tags that can be used for storing integrity information.
One more thing: where is the ECC RAM?
[1]: https://docs.kernel.org/admin-guide/device-mapper/dm-integri...
Also "Using gokrazy instead of Ubuntu Server would get rid of a lot of moving parts. The current blocker is that ZFS is not available on gokrazy. Unfortunately that’s not easy to change, in particular also from a licensing perspective."
I don't pretend to get how licensing works, but is OpenZFS and their licensing not an option here? I know its been really tricky with ZFS in general and I don't think its fully answered? but I'm not up to date with it.
Solutions such synology provide web interface, phone apps, and applications for syncing, photo management and backup. The units come with software for domain name, SSL certificates, monitoring, drive management, WoL, etc. There is a lot of software functionality useful for NAS built in.
You can get them new at the same price point only caveat that you need a motherboard that supports bifurcation.
Most of the resellers will answers questions or even flash firmware. Local Craigslist guys even offered to do it but they don’t know why you’re doing it, the eBay people understand.
[1]https://forums.servethehome.com/index.php?threads/firmware-e...
What, how? The only part of the chip that touches the ECC bytes should be the memory controller itself. The reads and writes going across the internal fabric should be exactly the same.
Also the PRO versions support ECC and seem to have the same iGPU.
On top of that, finding compatible ECC sticks is not easy and they can be quite pricey.
Using 64 gigs of this: https://media.kingston.com/pcn/PCN_KSM26ED8_16ME_C1.pdf
Given that was one of the requirements of the OP (integrated graphics), my point even though technically imprecise, still stands.
I have the same memory sticks. They cost £157.86 each 32Gb stick. That's over 2x the price of non-ECC sticks of otherwise similar specs.
Sounds like money well spent.
I wasn't making any value judgements though, only posting observations.
That said, it's much more likely that a faulty or low quality PSU will be the cause of that rather than the cable, and there's a lot already written out there about why a high quality PSU is important.
I replaced 2 of the 16GB ECC sticks in my main PC with 32GB ones a while ago and put them in my NAS, will probably do the same with the remaining 2 sticks to max out my RAM.
This shop [0] sells the same ones I have but it's even more expensive than it was last year, at £206.17 for 32Gb.
[0] https://harddiskdirect.co.uk/ksm26ed8-16me-kingston-technolo...
Also, just have enough RAM, but don't disable swap. Put swap on an enterprise SSD but let Linux handle it. It will do so cleverly. Whereas with RAM only you cannot use swap at all. Minor disadvantage is you should encrypt your swap. But with modern hardware that shouldn't be a large penalty.
1. A raspberry pi was picked as a single point failure point. This will be a headache within 24 months.
2. Daily power/thermal cycles might age some things more than a steady thermal load of always being on. Not a big deal and optimizing power consumption is totally reasonable, but a trade to track.
> When hardware breaks, I can get replacements from the local PC store the same day.
That wouldn't be my M.O. I don't want to go to a local PC store. The ones remaining are all either expensive, sell shit, or do both. I'd order from Amazon or whoever is cheapest according to a local price comparison website, and receive a replacement next day. It also costs less time having to do shopping (heck I even order groceries online, saving time). Besides, with warranty, it'd take longer than same day replacement.
Now, if the guy was running mission critical data solely I'd say 'OK', but nope:
> For over 10 years now, I run two self-built NAS (Network Storage) devices which serve media (currently via Jellyfin) and run daily backups of all my PCs and servers.
As such, it seems a silly requirement for the average content someone serves with Jellyfin.
I use 4x 4 TB HDDs in ZFS RAIDZ2 (running Ubuntu under Proxmox) with one enterprise grade SSD for OS and ZFS cache. All that Jellyfin data doesn't have to reside on any SSD. Not in this setup, and in most setups for home users: neither. You also don't require HA for such, to the point where even RAID is kind of silly.
I use restic for daily backups to a NAS in the same city, a Synology with RAID1 btrfs. But I don't back up any media served via Jellyfin. Seems silly, and I don't have fiber upload speed as of yet. The internet has backups of that. Usenet servers, for example.
None of these servers have a UPS, but in case a power cut the enterprise SSD (a Samsung NVMe costing more than 2x as much as a consumer version of same amount of storage) ensures the data is consistent. Both systems use FDE.
I also have an offsite, offline backup of our most important data. It is encrypted at rest but not I don't regularly rebase which is kind of stupid. But that is on me.
Either way, this fellow wants to have some kind of silent setup (for reasons) and yeah then SSDs make sense. But if you want value for GB/EUR or akin, even with the extremely low SSD prices of recent and even with the higher energy prices, then HDDs are still worth it. Especially for data as content served via Jellyfin.
I'm pondering about upgrading to more than 1 gbit for LAN but for now I don't see it is worth it. Also given my server (MicroServer 10 Plus) only has two PCIe and one is used for PCIe to NVMe and other for iLO I don't have the bandwidth available, a disadvantage of my current setup. I suppose I could use a USB to 2.5 gbit converter if USB could saturate it but then I'd use USB for high data transfers. It appears to me that because my switch is managed, pair bonding would be a better approach. Which the Synology also supports.
lately I've been experimenting with just putting the OS on spinning rust too. It seems like it's just so ingrained in us to put it on flash because of the order-of-magnitude gains in boot time, but for servers, is bootup time really a factor (damn supermicro motherboards take 1 minute+ post...)? Once your services are loaded to ram, how much does the OS actually hit the drives? If your OS supports ZFS-on-root, then you have one less point of failure.
Otherwise I can get a good quality SSD for $20. A hard drive for the OS is more money for worse performance.
If you're not okay with that, then $20 is cheaper than a hard drive.
"A hard drive for the OS" is referring to the latter situation.
I've had various consumer grade (but branded) USB sticks running Raspberry Pi OS, EdgeOS and Proxmox fail on me to the point where I use an USB to NVMe or USB to SATA. Why? I've tried industrial grade aMLC/MLC/SLC flash and performance is worse, and cost is high. For example, for Proxmox or EdgeOS you need at least 4 GB storage. They've never failed me but SATA or NVMe have high MTBF. On Raspberry Pis I've opted for log2ram on consumer grade (but branded) microSD with great effect. Another option is not log at all, or use e.g. rsyslog.