Building a fast all-SSD NAS on a budget
jeffgeerling.com
jeffgeerling.com
Love seeing these projects from him, but this is a rare miss in my opinion. This is the strange middle ground where it's closer to professional than "budget", by quite a bit. At some point, only a tiny tiny fraction of users need 40TB of space. I guess what I'm saying is, this isn't so much a NAS as a specialized youtuber video recording appliance. We have a number of home-hosting users on our platform that run entire racks filled to the brim - and while that's awesome - it's just not what "budget" or "NAS" implies, and it's an extremely limited audience.
Agree though that this was a bit of a miss; the cost of drives here dwarfs everything else, and it makes the choice of case, MB, etc. nearly irrelevant and needlessly cost optimized. Why not spend an extra $60 and get a 5.25->6x 2.5 hot swap cage, for example?
I'm also trying to get my homelab set up to be a bit more robust / automated / reproducible, and trying out a bunch of different ideas. I'll likely settle on something else in a year's time, but the main motivation for now was to get all the storage into my rack on my 10 Gbps network.
At one point when I was younger, that was all fun. Now, I just want it to work so I can go back to having a life and let others do all the experimental stuff. But now I get to live vicariously through people like you posting their results. Good stuff!
The only other thing that annoyed me is calling TrueNAS Core (FreeBSD) a "Distro" in the same sentence as a Linux distro.
The reason to question this is that 40TB seems small if you want to have a NAS for small video editing studios. And for personal use, you probably not going to need more than 2TB work set paged in at any given moment.
This will be considerably faster for working with "immediate" needs of video files rather than over a 10GbE network.
like, a difference of 900MBps over network vs 2500MBps with local sequential read/writes on NVME SSD on same motherboard.
HDDs can't provide the latency that I need, unless I was running a dozen or two to overcome the horrible seek speeds (or had at least double the RAM so the entire project would be cached in RAM).
Especially if the NAS was hyperconverged-in-the-small (i.e. had the spare compute to render out proxy footage for uploaded files in the background.)
Storage needs for any pro video workflow get very large, very quick.
[1] https://diskprices.com/?locale=us&condition=new&capacity=12-...
https://www.truenas.com/docs/references/slog/
The short version is:
- ZFS caches reads into RAM. Gobs of RAM helps. You can also add a secondary read cache device (L2ARC).
- ZFS can caches synchronous writes to an SLOG device, but if you've disabled synchronous writes it won't make a difference. An SLOG device makes no difference for asynchronous writes.
I use ZFS for my Time Machine backups (among other things) and have synchronous writes disabled for the Time Machine datasets.
Before we had ZFS, the traditional way to speed up NAS was a raid controller with a battery-backed RAM cache.
https://www.reddit.com/r/MacOS/comments/lh0yjc/configure_a_t...
Basically it creates a dataset and a Samba share, and the configures Samba with a couple plugins so that: 1) a dataset is created for each user that connects to the share, so there's per-user datasets; 2) a ZFS snapshot is created automatically whenever Time Machine disconnects from the share. (2) is so that you can roll back if need be.
In the past, I've had Time Machine decide that the sparse image it uses for a backup is corrupt and it wants to zero it out and start over. The ZFS snapshots lets you recover from that, though I haven't had it occur under macOS Monterey and/or since I switched to SMB from AFP.
On the parent dataset I disabled sync.
If you google for "TrueNAS Time Machine" you'll find some discussions of all this in the TrueNAS forums.
I suspect the "Purpose: Multi-user Time Machine" selection is doing a lot of heavy lifting, but it's hidden away under a shiny frontend.
The "Multi-user Time Machine" is doing a bit of heavy-lifting yes. If you're really interested I can pull-out the details of the Samba config and how the ZFS dataset is configured. I do think it's using some TrueNAS specific Samba plugins, but I guess in theory those are open-source and could be compiled for Debian.
> If you're really interested I can pull-out the details of the Samba config and how the ZFS dataset is configured.
I would be thrilled to try implementing that, yes!
So is mine. but good golly, who the heck wants to keep doing their day job at home? For my home gear, I try to keep the IT stuff to a minimum. I've been using FreeNAS/TrueNAS for maybe a decade now and it just chugs along, no muss, no fuss. Migration and upgrades have been trivial.
But anyway...
Here's my samba config:
https://gist.github.com/jaysoffian/43636b535ec37fdd31f0514bc...
The "ixnas:" options require the Samba ixnas VFS. It's here, but I actually don't see these options in it:
https://github.com/truenas/samba/blob/release/22.02.3/source...
Here's a discussion of what they do though it's pretty obvious from their names:
https://www.truenas.com/community/threads/configuration-opti...
The fruit VFS is part of the normal Samba distribution.
Here's the dataset properties:
https://gist.github.com/jaysoffian/2d2ac21d60bc40194e01f362c...
Good luck.
There's a reason that a big difference in price exists between a quad-level-cell 2TB SSD and an expensive enterprise grade one with a much higher TB-write-before-dead rating.
This might look cool but check back in a few years and see how much of the drives' cumulative write lifespan is worn out.
I also cannot even imagine spending $4000+ on a home file server/NAS with copper only 10GbE NIC and it not having at least one 10G SFP+ interface network card.
Okay, so he wants it to be tiny? But in a home environment the major problem is more power consumption and noise, so you can often go with a well ventilated 4U height rackmount case for full size ATX motherboard, which is roughly the size of a midtower PC case turned on its side.
This lets you use motherboards that will have enough PCI-E 3.0 x8 slots for at least one dual-port Intel SFP+ 10G NIC which are very, very cheap on ebay these days.
I had assumed this is what he was using the TLC SSD for. If that’s the case, so long as there isn’t much writing to it, it should be fine.
How many full-drive writes does your average video editing server need? I would expect a pretty small number. The average source file is sitting there for weeks or months.
For a use case like database transactions, log storage and frequent data dumps, the game changes quite a bit. I would definitely shy away from the QVO drives for that use case.
I've had these drives in service for about 8 months in my regular NAS, before transferring them to this new build, and they are all checking out okay still.
But this is also why I'm doing the striped mirror plus a hot spare. The only real challenge would be if the drives have a firmware issue, and they all die at the same moment after like 4 years due to a bug (like how HN's servers died...).
I don't know what you're using your NAS for, but the author is using it as scratch space for raw video files. It's not an OLTP DBMS or anything. It just needs really fast ingest of files beyond the capacity that a DRAM cache can provide.
> I also cannot even imagine spending $4000+ on a home file server/NAS with copper only 10GbE NIC and it not having at least one 10G SFP+ interface network card.
The author's editing rig doesn't necessarily have a 10G NIC, let alone being attached to a 10G switch with runs of CAT6a; and there's only one device talking to this NAS at a time (as the author cannot be in two places at once.) So what'd be the point?
But copper is usually a little simpler for consumer/prosumer devices. Someone does make a Thunderbolt to SFP+ adapter but that things like $300!
At 500x whole-drive rewrites I think I got my money's worth out of an $84 1TB drive.
Reports of SSDs write span issues have been greatly exaggerated ;)
At least nowadays even with QVOs thats not something consumers have to think of much anymore.
These SSDs have 8TB, so to exceed its write endurance Jeff would need to write 4TB to all of them each day for 3 full years.
That doesn't make sense. The author is not using raidz. In ZFS terminology, he's mirroring two striped vdevs (like raid10), plus using a hotspare. And that is a bad choice, as he gets only 16 TB usable, and the mirrored stripes could fail with certain combinations of two drives failures.
Instead, he should have set up a raidz2 across 5 drives: 50% more usable space (24 TB), can tolerate any two drive failing, and it would give him higher performance on sequential I/O.
I have a raidz2 pool across 6 spinning rust 18TB hard drives, and my server can handle 900 MB/s sequential reads, and 700 MB/s sequential writes (benchmarked locally). It it was built on 5 SAMSUNG 870 QVO SATA III SSD drives like the author, it would certainly get to 1.5+ GB/s in both sequential reads and writes. In other words the bottleneck would shift to the 10 Gb/s network (1.25 GB/s). For comparison the author discloses his writes are limited to 700 MB/s (over the network, so his bottleneck is not the network link but local I/O contention).
That's right, a 6-HDD raidz2 matches the write performance of his 5-SSD raid10 setup. And that's because the write overhead of raid10 is 100%, while the write overhead of a 6-drive raidz2 is only 50%.
Edit: Oh and the strange performance issue that he noticed "disappeared" after a while is most likely due to cache recovery. The drive can sustain writes of about 490 MB/s for a little while then it drops to 170 MB/s, but after idling 5 minutes it can recover the initial speed. See https://www.tomshardware.com/reviews/samsung-870-qvo-sata-ss...: "As we noticed with the 1TB model, the 8TB model’s cache recovery mechanisms work similarly. After letting the drive rest at idle for 30 seconds, the 870 QVO gains back 6GB of its cache. It recovers fully with 5 minutes of idle time. "
One concern Wendell had with RAIDZ2 was the potential for the older Xeon to be a bottleneck for writes. It probably wouldn't be, but I didn't do too much testing with that layout. I might still, we'll see.
Best bet is to pick up an old HP Z620 of find someone who is upgrading their old Xeon homelab. Generally its a choice of cheap, quiet, energy efficient and you can only pick two of these options.
Someday I'll have a proper closet but right now my rack is under the kitchen and my wife complains when the servers get above about 60 dB.
I actually built my desktop as a near silent PC with the goal of it becoming my NAS when I upgrade but I didn’t understand the implications of running ZFS without ECC at the time.
So while I'd certainly sleep sounder with ECC on it, it's not a complete horror show like someone likes to portray it as.
I have had one stick give me issues over the last 10 years but that is out of probably 20 sticks.
For a video editor, at least one that's been around long enough to remember DAS and SAN solutions, $4300k for 40TB of edit capable storage is cheap.
Perspective is everything.
And peeps wonder why we think $4500 is cheap!
Or is it one of those things that would be cool to have but not really necessary?
For installing outside of a machine room/closet/center, if you're using 2U of height, you might also fit a PSU with a larger and quieter fan, since all the Flex PSUs I've had come with noticeably loud fans. (I replace them with Noctuas, but it isn't a fun kind of soldering, IMHO.)
The components from the build would also fit in a Supermicro 1U short-depth chassis, especially if you can go a little deeper in your cabinet. (My new K8s server got a used Supermicro 1U chassis for ~$60 shipped, including a PSU. In the photo on https://www.neilvandyke.org/kubernetes/ , it's the 1U immediately below the 4U.)
If your IO all goes over a 10 Gb link, that will be your bottleneck before ZFS is.
Since 1 MB record size was used, this may be a candidate for a bunch of mirrored HDDs with a few TB of nvme for l2arc and ZIL. But with a bunch of SSDs already on hand, why bother?
Each 980 Pro can do sequential reads at 7 GB/s. ZFS with record size 128 KiB will do IOs at a size that will perform about as well as sequential regardless of the actual pattern.
AFAICT ZFS has no need to push your drives beyond 10% of their overall capability assuming the clients are performing sequential reads or the reads are at least 128k. A more stressful read test would be a 4k random read where the working set doesn’t fit in the ARC. This would trigger a lot of read inflation.
The write side is a bit more complicated as the drives may be able to write at over 5 GB/s if they have 30+% of their NAND erased, else the write rate will be closer to 1.4 GB/s. With a mirror this is pretty close to the network bandwidth. With some reads mixed in for COW the drive could become a bottleneck before the network or ZFS.
Do you really need to save all of that footage? I would think keeping the pro res footage for the current projects on the workstation and reencoded archive video on a NAS would be sufficient. I'm not a video professional, but I suspect it's easy to fall in the trap of thinking that you need to save everything in highest possible quality in case you need it later, but what are the realistic chances of that? If you end up needing some old footage again, AV1 coded 4K or even HEVC 1080p would probably be just fine. The final result are Youtube videos after all.
I know he mentions editing from it, but that's enough space for more than a week of pro res video.
On the flip side, I think it would help a lot a lot a lot if there was more discourse out there about compressing high-dynamic-range content. AV1 supposedly has some capability to do a good-ish job with HDR. The idea of taking reels of raw video and spending a couple days squishing it to 1/10th the sizes but preserving the quality/flexibility very very well is a value proposition that I dont think is clearly attainable, even though it seems technically perhaps within reach.
Right now, I think the general feeling is, raw is raw & everything else bakes in a vast amount of assumptions & constrictions. Some advocacy that compressed video can be as flexible, as dynamic, as capable need to be more present, elaborated, & proven before anyone's going to be comfortable throwing away the bits. These are people's life's works & a couple hundred or thousand dollars a year more in storage isnt a real factor for such integral, near work.
Have you seen Jeff's videos? They're mostly him talking to the camera in what I think is his basement, or closeups of various PCBs. If you're running a stock photo company I get wanting to keep footage at max quality, but once your Youtube video is published, do you really need to save all those alternate takes and cut out bits? I just don't see them being very useful in the future, and if you do end up needing some of it for B roll or whatever, would a couple of seconds of compressed video really make a drastic difference for the project as a whole? Especially since the final result ends up on Youtube and consumed on a phone screen or a TV two meters away.
Much of the content that's produced today is not trying to be timeless classics with endless rewatch value, it's ephemeral and only really relevant for a short time. If you can reupload your videos to other video hosts in the future that's probably good enough. Nobody is watching reviews of three year old Raspberry Pi add on boards for the low noise in dark parts of the image, or the accurate color reproduction, or even the 4K instead of 1080p resolution. Kill your darlings!
Why comment your code or name your variables well if the project will be over soon? No one is buying your product for the clean code behind it. If you need it later just use a reverse compiler. It’s good enough for a project you may never need!
People do things for the art and to be proud of the quality. And you don’t want to throw that out!
I am a video professional, and you always keep your originals. However, the part being left out of the conversation is that as a storage pool for editing goes, this isn't deep storage. Content is typically only left on the edit storage while the project is active. Even the TFA mentions he goes all the way to cloud cold storage.
Edit storage has always been expensive, and maintaining capacity was always a juggling act. Just because terms like NAS are being used, one should not think of this type of storage as a dump it and forget it type of storage.
There are many levels of professional. On the high end, the footage from the camera is copied to multiple hard drives on set. These are the backups, and us old timers still use the terms camera originals. These are as sacrosanct as film negatives, only there's magically multiple copies. This data gets transferred to the editor's storage. Once the edit is completed, the edit session and other content used in the edit may be transferred to the camera original drives for archival. On the other end of professional, you take the SD card out of the camera and transfer directly to the edit storage. In these situations, woe be unto thee that doesn't make proper backups. At that point, it's really more pro-sumer than professional, but hey, if they're getting paid, they're professional enough.
And yes, I did see the part about editing, but are you really editing nine days worth of footage at once? Would it make more sense to put less but faster storage directly in the workstation?
Doing that kind of editing does require access to all of that footage at any moment. If you're doing supervised sessions with the client sitting next to you means that they on a whim could decide they want to work on a totally different part of the content. When the client is footing the bill by the hour, you don't get to waste time "loading" content.
It's really one of those things that until you've walked a mile in another person's shoes, it's best to not go blindly suggesting "better" based on one's limited knowledge. Asking questions for a better understanding is a totally acceptable way of learning, and I'm all for it. We're danger close to the former. Hopefully, I'm not sounding like an asshat leaning towards the latter.
And I repeat another point. Using his own numbers, this editing NAS has space for over 220 hours, more than 9 consecutive days. What kind of music video production even has that much video footage? Like I wrote in my initial comment, I fully understand using plenty of uncompressed video while editing, it's the saving of hundreds of GB per finished and delivered Youtube video I question.
Some videos are 10 minutes, others are more like 20. On average I end up with 100-200 GB/video by the end.
So you're correct that I don't need 14TB (that's overkill) for now... but anyone who plans storage knows you're better front loading the extra capacity if you can afford it, because needs always seem to grow faster as time goes on (e.g. 4K to 8K, ProRes LT to something with more color space maybe...).
And right now I keep all my "A-roll" (scripted takes, usually with about double the run time of the actual video), but I'm considering setting up a script to automatically convert those clips to H.265 afterwards so I can still have reasonably good archival footage but at like 1/100th the file size.
All B-roll shots (Dolly shots, handheld closeups, illustrations, etc.) I save at full original resolution, because they are often useful in follow-up videos. But that's usually only like 30-40% of that 100-200 GB per project.
(Edited to add: sometimes there's a unicorn video that also has raw Timelapse footage where you roll at normal rate then generate a Timelapse afterwards (and use some of the real-time footage). Those can have hours of footage and a couple have gotten beyond 400 GB. But I try to limit that just because of the hassle.
Easier to have one camera devoted to the Timelapse, and another for some other shots for real-time use.)
Please, don't cheapen my career shooting actual timelapse! ;P
I usually like to use my Nikon D700 for that.
FFmpeg has gotten decent AV1 support with SVT-AV1 recently. I think it would make for an interesting video to compare different codecs and see which one is "best" for a certain set of trade offs. You could make a script that encodes a video snippet while varying parameters like codec, preset, crf, grain, 10-bit, hardware or software encoding, etc, use one of the automatic quality metrics like Netflix's VMAF, and plot size, quality, encode time, power usage against each other.
This creator can't do other editing gigs besides his own content? Granted, he's probably pretty busy keeping up with his own stuff, but don't put baby in the corner by suggesting they only work on YT content. You'd be amazed at how many balls an editor can juggle. I currently have a couple of corporate gigs in various stages, wrapping up a music video, and a plethora of personal side projects on my system right now. Currently in preproduction meetings with new client to specifically help them start their YT and other social media content creation. So because the conversation started with a YT content creator means the conversation can't be expanded to introduce you to the bigger picture of the same topic?
Don’t take this the wrong way, I’m genuinely curious… what is a video professional doing on HN? Was it a past life? Is tech just a side interest? I see many non-tech professionals on here (doctors lawyers too) and I’m always curious what drew people in!
probably more than you really wanted to know, but that's it summed up in graph. i've learned that if a job requires working with a computer in any way, knowing how to bang code will come in handy even if it's not a coding job. just knowing how to automate a few things here and there makes cubicle life less mundane.
Ubiquiti have a cheap fiber optic switch you could try. You could also try a 2x 10G SFP+ configuration, which would give you 20 Gbps (but only 10Gbps per client).
Maybe not THE BEST (tm) choices. But I was getting bad decision paralysis choosing parts.
Very entertaining and I learn a ton.
I have a Raspberry Pi NAS with 10TB (2TB of redundancy) but it’s in a box made out of MDF. And while I had fun making it, let’s just say I’m not a woodworker.
I’m amazed at how professional the stuff he does looks all while making it seem easy.
~ cat /etc/auto_master
#
# Automounter master map
#
+auto_master # Use directory service
#/net -hosts -nobrowse,hidefromfinder,nosuid
/home auto_home -nobrowse,hidefromfinder
/Network/Servers -fstab
/- -static
/- auto_nfs -nobrowse,nosuid
~ cat /etc/auto_nfs
/System/Volumes/Data/Users/ahepp/nfs/public -fstype=nfs,vers=4,resvport,rw nfs://10.128.1.2:/zpa/nfs/public
I haven't done a lot of performance testing, but I have no problem playing 10MB/s blurays off the NFS share (and that's via wifi)On the FreeBSD server in the closet:
root@tlon:~ # cat /etc/exports
V4: /
/zpa/nfs/public -mapall=nobodyWhat's really confusing me is the 10G lan configuration because it's leaving little headroom above uncompressed 4K 4:4:4 10bit (12bit is broadcast and where source permits archive standard). What about multiple streams for a/b'ing a grade or first/ second camera edit roll?
Ed. added "really" replaced "for" with "above"