ZFS 2.1.6
github.com
github.com
There is no need for FreeNAS or any of that stuff, I don't use all those features. Just run like 3 commands to create a zpool, install samba service and that's all I need from a NAS. Cron job to run a scrub every month.
If ZFS file system made a sound, it'd make a satisfying and reassuring "CLUNK!" sound.
The common tasks are made easy: snapshot scheduling, snapshot replication, Samba share with a few clicks, backup to cloud, users management, one click install of applications, SMART tests, etc. Config management can be a headache.
The only other thing, you may be forgetting, is an app you never knew you needed.[0] ;)
[0, an interactive, file-level Time Machine-like tool for ZFS]: https://github.com/kimono-koans/httm
I'm in the same boat, all I have on my NAS is Samba, zrepl (ZFS snapshots and backups) and node exporter (monitoring agent for Prometheus - handles SMART, etc). It's running Arch, and I've never had my setup "break" in more than 10 years of using this distro (though this particular NAS is not that old).
I don't care for a DB, web server and god forbid random applications doing who knows what on my NAS. But if you want your NAS to do everything, it should be great. To each their own, I guess.
Also, TrueNAS is FreeBSD which may or may not work on dinkier hardware. I'm specifically thinking of random watchdog problems on some older Realtek Gigabit NICs, where you had to use a specific patched driver which, at one point, didn't support the latest version.
Plus the setup now gets intricately coupled with the host OS. That's a big no for me. I want to treat the host OS as disposable, ephemeral and decoupled for peace of mind. The host's only job is to serve files over samba protocol and run scrub cronjobs. That's it. It gets 2 vCPUs and 24GB of ram. Stays put.
But then, the whole thing should be marketed more like "Proxmox with batteries included".
It's not though. Even if you make all the "NAS software" part of the jails (or the half-baked Kubernetes in TrueNAS SCALE), it's not even close to Proxmox when it comes to managing VMs, containers/jails, firewall, networking, etc.
Though NFS can only be shared from the global zone, so if you need NFS you don't get to play with zones.
I've been a ZFS user for maybe 6 years. I had data loss this year.
It was quite interesting. Writes would fail on multiple disks at the same time. That's why it was data loss. A normal disk failure wouldn't look like that, failures should be spread out and redundancy would help.
It turned out to be a bad power supply. I was able to predictably get it to corrupt writes with the old PSU. Then I replaced the PSU. No longer failed after that.
I wouldn't have guessed to suspect the PSU. It was a frustrating experience but in the end, ZFS did help me detect it. On ext4 or ffs I wouldn't have even been aware it happened, let alone could confirm a fix.
How does power fluctuations or cutoff cause data loss in ZFS?
I frequently force shutdown my laptop and PCs with ext4 when they hang, and never lost data.
How do you know?
I've seen a transient error due to a misconfiguration of ALPM that XFS/ext4/etc probably would have never caught, but ZFS did.
And FYI, an power supply failure does not manifest the same as a power outage.
I also experienced some random reboots, things that also looked to me like bad memory... I suspected bad memory at some point. But swapping the PSU did the trick.
A large file copy would predictably trigger it. Other times it was random.
I swapped every server part before the PSU. Updated Linux, downgraded, tried kernel options... At some point I thought btrfs was the problem so I created an mdadm ext4 raid but got the same problem
It was a btrfs raid 1, lots of fs errors, and files missing until restart, as some disks were down. But didn't lost any data (take that ZFS!) besides the one being transferred as the disks/array went down anyway
This doesn’t sound like data loss but rather a fault preventing writing… it would be data loss if confirmed writes were lost.
You can disable write caches for safety, but note that this is very hard on performance.
POSIX systems are pretty lax with this sort of failure. write(2) and close(2) can succeed if you write to cache. If the actual write failure occurs later there is typically no way to let your process know.
After replacing the SSD twice and then rebuilding everything using a different motherboard, it turned out it was the PSU.
Really hard to diagnose problems like this.
This is true for every mature/production file system. You can do the same with mdadm, ext4, xfs, btrfs, etc. The only constraint would be with versions in that it would be a one way street. Can't necessarily go from something new to something old, but other way round is fine.
On the other hand, if you move drives from a hardware raid and put in new drives, some (all?) controllers will read the raid config from memory and offer to build the same raid on the new drives. That's completely-out-of-band. Depending on the controller, even changing the order the disks are plugged in can give you weird results.
[1]: https://openzfs.github.io/openzfs-docs/man/8/zfs-set.8.html?...
And yet, this thread right here has more and better info than my searches ever turned up, which were mostly reddit and stackoverflow posts and such that somehow managed never to answer the question or had bad answers.
The one complaint is that I found that to be true for almost everything with ZFS. You can read the manual and figure out which sequence of commands you need eventually, but "I want to do this thing that has to be extremely common, what's the usually procedure, considering that ZFS operations are often multi-stage and things can go very badly if you mess it up?" is weirdly hard to find reliable, accurate, and complete info on with a search.
The result was that I was and am afraid to touch ZFS now that I have it working and dread having to track down info because it's always a pain, but I also don't really want to become a ZFS wizard by deeply-reading all the docs just so I can do some extremely basic things (mirror, expand pools with new drives, replace bad mirrored disks, move the disks to a new machine if this one breaks... that's about it beyond "create the fs and mount it") with it on one machine at home.
The initial setup reminded me of Git, In a bad way. "You want to do this thing that almost every single person using this needs to do? Run these eight commands, zero of which look like they do the thing you want, in exactly this order".
I'm happy with ZFS but dread needing to modify its config.
> I also don't really want to become a ZFS wizard
Admittedly, ZFS on Linux may require some additional work simply because its not an upstream filesystem, but, once you're over that hump, ZFS feels like it lowers the mental burden of what to do with my filesystems?
I think the issue may be ZFS has some inherent new complexity that certain other filesystems don't have? But I'm not sure we can expect a paradigm shifting filesystem to work exactly like we've been used to, especially when it was originally developed on a different platform? It kinda sounds like you weren't used to a filesystem that does all these things? And may not have wanted any additional complexity?
And, I'd say, that happens to everyone? For example, I wanted to port an app I wrote for ZFS to btrfs[0]. At the time, it felt like such an unholy pain. With some distance, I see it was just a different way of doing things. Very few btrfs decisions with which I had intimate experience, do I now look back on and say "That's just goofy!" It's more -- that's not the choice I would have made, in light of ZFS, etc., but it's not an absurd choice?
> "what's the procedure if your motherboard dies and you need to migrate your disks to a new machine?"
If you're setup is anything like mine, I'm pretty certain you can just boot the root pool? Linux will take care of the rest? The reason you may not find an answer is because the answer is pretty similar to other filesystems?
If you have problems, rescue via a live CD[1]. Rescuing a ZFS root pool that won't boot is no joke sysadmin work (redirect all the zpool mounts, mount --bind all the other junk, and create a chroot env, do more magic...). For people, perhaps like you, that don't want the hassle, maybe it is easier elsewhere? But -- good luck!
[0]: https://github.com/kimono-koans/httm [1]: https://openzfs.github.io/openzfs-docs/Getting%20Started/Ubu...
Here for this. Already delighted and amused. ;)
> edge of my seat the whole time.
In the future, you may want to try creating a sandbox for yourself to try things? I did all my testing of my app re: btrfs with zvols similarly:
sudo zfs create -V 1G rpool/test1
sudo zfs create -V 1G rpool/test2
sudo zpool create testpool mirror /dev/zvol/rpool/test1 /dev/zvol/rpool/test2
sudo zfs create -V 2G rpool/test3
sudo zfs create -V 2G rpool/test4
sudo zpool replace testpool test1 /dev/zvol/rpool/test3
sudo zpool replace testpool test2 /dev/zvol/rpool/test4
sudo zpool set autoexpand=on testpool
...> Here for this. Already delighted and amused. ;)
Haha... yeah, I didn't intend that as a brag or badge of honor or anything—more like a badge of idiocy—but you don't play Human Install Script and a-package-upgrade-broke-my-whole-system troubleshooter for several years without learning how things fit together and getting pretty comfortable with system config a level or two below what a lot of Linux users ever dig into. Just meant I'm a little past "complete newbie" so that's not the trouble. :-)
> In the future, you may want to try creating a sandbox for yourself to try things? I did all my testing of my app re: btrfs with zvols similarly:
Really good advice, thanks. I was aware it had substantial capabilities to work in this manner, but using it this way hadn't occurred to me. Gotta get over being stuck in "filesystems operate on disks or partitions on disks that are recorded in such a way that any tools and filesystem, not just a particular one, can understand and work with" mode. I mean I'm comfortable enough with files as virtual disks, but having a specific FS tools, rather than a set of general tools, transparently manage those for me, too, seems... spooky and data-lossy. Which I know it isn't, but it makes the hair on my neck stand up anyway. Maybe my "lock-in" warning sensors are tuned too sensitive.
Now to figure out how to run those commands as a user that doesn't have the ability to destroy any of the real pools... ideally without having to make a whole VM for it, or set up ZFS on a second machine, and—initial search results suggest this may be a problem, for the specific case of want unprivileged users to run zfs-create without granting them too much access—on Linux, not FreeBSD :-/
And to be fair, I think ZFS could be better in this regard. Some commands can put your pool into a very sub-optimal state, and ZFS doesn't warn about this when you enter those commands. Heck even the destroy pool command doesn't flinch if by chance nothing is mounted (which it may well be after recovery on a new system).
I found it helped to watch some of the videos from the OpenZFS conferences that explains the history of ZFS and how the architecture works, like the OpenZFS basics[1] one.
But I agree that the documentation[2] could have a lot more introductory material, to help those who aren't familiar with it.
That said, I echo the suggestion to try it out using file vdev's. For larger changes I do spin up a VM just to make sure. For example, it's possible to mess up replacing a disk by adding new disk as a new single vdev rather than replacing the failing one one, so if I feel unsure about it I take 15 minutes in a VM and write down the steps.
Again, this is something I feel they could improve. Adding a single-disk vdev to a mirrored or raid'ed pool should come with a warning requiring confirmation.
On the bright side, I've been running my pool since 2009, and have never lost data despite a few disk failures and countless unexpected power-outages without PSU. And I just run it on consumer hardware without ECC because that's what I got. Been up to 8 disks, now down to 6 and will soon go down to 4 once the new disks arrive. Send/recive ensures the data is just as it ever was on the new configuration.
I would encourage anyone that does not have an system admin background to use FreeNAS(truenas now) or something
Their are default security and maintenance things (like periodic scrubs of the pool) that these nas operating systems set up by default
??
The OS needs to implement zfs, like any other filesystem/host combination. Any filesystem can be "moved" to another host simply by attaching the drives, assuming you have hardware level compat and filesystem level compat.
OK the exception is when you have host-based hardware level encryption (ie, key in TPM or other security chip). In that case, zfs and otherwise, you can't just move the drives.
Fedora + BTRFS + Deja dup daily backups have been my default for months now, but I am tempted to give it a go to perhaps Ubuntu JJ + ZFS + Deja dup. Not that I have any issues with my current setup. It's just for the sake of, you know, try new things.
I'm not familiar with Deja dup, but the usual method of backing up a ZFS filesystem is to snapshot it and then stream the snapshot (just a file, can be compressed) into the remote store of your choice.
The usual way is ZFS send and receive to another ZFS server. But almost no cloud storage provider supports that, other than rsync.net (and zfs.rent), which starts with $60/month. So you have to set up and manage another ZFS server with secure remote access, PIA if you ask me.
For a while now, my backup process has been essentially to run
zfs send -L -c -P "${snap}" | zstd - > "${out}"
gsutil -o GSUtil:parallel_composite_upload_threshold=150M cp "${out}" "${gcs_path}/"
It seems to work well. Am I carrying some unrecognised risk that I wouldn't be carrying if I were backing up into another ZFS filesystem? How does that risk balance against the risk of data loss on the backup server? i.e. I am pretty sure GCS won't lose my data, but not quite so sure about the other services you mention.Just chiming in here to say that I endorse every aspect of the answer that 7e has given you.
The error occurs in transit or on client side, not necessarily on remote.
Also managing a chain of incrementals can get unwieldy.
Ask r/zfs.
Snapshots are incremental, so management of your retention period is an active process and may require combining older snapshots together when you drop them.
With a zfs server on the other end you can zfs send / recv the other way to restore, and it correctly manages deleted snapshots.
They are but, unless the -i flag is specified (it isn't in my example above) then zfs send will send a complete replication stream, not an incremental.
I honestly use ext4 for everything based on ignorance and familiarity.
Checksums if the mobile hardware is damaged, snapshots if you delete something in your daily driver, native encryption if laptop is stolen, backup of work data via ZFS send if you have a ZFS server, compression to use limited disk space in laptops, errors can be fixed with copies=2, etc.
- I can easily create new subvolumes (filesystems or block devices), for different mount points, containers, other distributions, VM disks...
- The free space is shared between all of them, no need to anticipate a partitioning layout
- I have a service that automatically snapshots some of them, others not
- I love the ZFS tools, I'm very happy to use them on all my machines
- I use NixOS everywhere and so it's also easier to share the same conf on all my machines
- I can send volumes easily from one machine to the other (zfs send)
I probably miss other things, but I also love consistency ;)
- no overlayfs yet (https://github.com/openzfs/zfs/pull/9414) - the docker zfs driver is very slow and buggy - best to create a sparse zvol with ext4/xfs/btrfs and use this for /var/lib/docker
- no directio yet (https://github.com/openzfs/zfs/pull/10018) - not sure if this is so useful beyond specialized databases that utilize that.
- no idmapped mounts yet (https://github.com/openzfs/zfs/pull/13671) - this would be useful for lxc/lxd
- async dmu (https://github.com/openzfs/zfs/pull/12166) - complicated patchset would transform a lot of operations into callbacks and probably increase performance for a lot of workloads.
- namespace delegation was recently merged: https://github.com/openzfs/zfs/commit/4ed5e25074ffec266df385... with idmapped mounts it should be possible to have zfs datasets / snapshots inside a unprivileged lxc-container which could be super cool for lightweight container things.
caveats:
- no swap on zvol at the moment (https://github.com/openzfs/zfs/issues/7734) - this is some hairy memory allocation problem beneath it and I don't really use swap - just something to be aware off.
- arc vs. pagecache / performance (https://github.com/openzfs/zfs/issues/13736) - zfs is usally very fast but there are some edge-cases where it's not at the moment.
That's great. I had a conversation with the owner of rsync.net about why they require zfs-enabled customers run on a separate bhyve VM (thus requiring a large minimum). Iirc it boils down to zfs destroy permissions somehow affecting other tenants. I see mention of zfs destroy in the linked patch, I'm curious how that relates to rsync's experience and requirements.
On the other hand it's not that hard to export and re-import a ZFS file system from a workstation or laptop to defragment manually - it's harder to do that with a NAS with lots of drives when you don't have a second set of drives just hanging around.
https://openzfs.github.io/openzfs-docs/Basic%20Concepts/dRAI...
Expand your ZFS array over a network. It's not expansion, BUT it can be used to grow a pool size.
And a hard disk array with 4 drives in raidz (like raid 5).
Main advantages: * I moved my root drive to another drive and did it entirely online without needing to cp or dd everything just add new drive to pool, remove old drive (and wait until it's done, but you can keep using it while it copies) * It's way faster than MD, md spends 10s of hours initializing or resilivering big drives. zfs just formats it and your good to go. * Subvolumes and snapshops are great. You can give a folder different file system properties. (Eg make /var/log compressed and disable it for /var/cache/pacman because it's already compressed) * Docker is quicker. * Although BTRFS has snapshots I do find ZFS's more intuitive. (eg subvolume can be nomount in ZFS which allows you to organize and apply options recursively easier)
Check for bugs in Ubuntu ZFS before installing updating. They do it differently than other OpenZFS options.
I'm not currently using ZFS send for backups. So I can't compare it to Deja dup.
I love it but two problem I ran into on Ubuntu:
1. Docker failed bcs of the custom implementation in Ubuntu (Don't remeber the specifics).
2. I repeatedly run into the problem that rpool (root) or bpool are supposedly full and automatic snapshots break.
I hope those problems get fixed soon.
* The cpu usage can be higher if you use compression.
* I used to setup the main pool with SSD cache with HDD data store. It's okay for most use case except slow first time open of applications. But for demanding games it usually has very bad performance.
* ZFS also manage cache by itself so it needs some memory. Personally I'd like to set a cap for it, because otherwise when I open some memory hungry program it tries to write cache onto disk which sometimes makes the system slow. This happens when I only use SSD as cache. Not sure if it can be better with all SSD solution.
* I ran into problem with docker as well. But after change some configurations it works perfectly.
* I use zfs snapshot and zfs send/recv for backups and it is really awesome. I've written a blog about my personal offsite backup solution that you may find useful: https://www.binwang.me/2021-09-19-Personal-ZFS-Offsite-Onlin...
Separately, at work I managed to persuade them to configure our two new servers with ZFS, each with 168 drives for a total (after RAID) of 1.5PB each. Works like a charm, and is quite happy to saturate a 10Gb ethernet connection when accessing from another server, with transparent compression turned on. I would not hesitate to recommend ZFS to anyone.
The scripts in that post are a little outdated, so if you would like to see updated ones, contact me [2].
The end result is a setup that:
* Has parallel upload to saturate my slow upload connection.
* Allows resume with `make -j<cores>`
* Allows me to backup things that are only on my server, such as my Gitea instance [3], to my desktop, thus having my server and desktop serve as backups to each other.
* Allows me to delete snapshots older than a certain number of snapshots (set to 60).
* Has automatic loading and unloading of encryption keys, to reduce the exposure time. (However, I wish raw sends on encrypted datasets worked. But it is nice that I can load the keys and download individual files instead, which I've already had to do.)
Perhaps I should add an update post to [1]. If there's enough interest, I will.
Edit: I should make clear that my ZFS is only my home directory. I do not use it as a root or boot filesystem because it runs off of a mirror of hard drives, and my root/boot drive is an SSD and can be destroyed without pain because all of my configs are in a repo in my home directory. Thus, it would be trivial to rebuild my system.
[1]: https://gavinhoward.com/2021/02/adventures-in-backing-up-dat...
No real complaints except that it complicates updating Arch: kernel updates usually have to be delayed 24-48 hours while the ZFS packages get rebuilt, or in some cases patched upstream. syncoid/sanoid also poses a problem because it relies on a few Perl packages from the AUR that are a pain to rebuild.
I know the Arch devs loathe people doing partial updates, but occasionally delaying a kernel update for the sake of ZFS has yet to bite me in the ass. I would strongly advise against DKMS, though. Use a binary module or build the package yourself. The majority of problems I had with Arch+ZFS boiled down to misuse of zfs-dkms.
zfs-dkms + "linux-lts" (longterm kernel) seems to work better than zfs-dkms + linux (stable kernel) if you're intending to do frequent system updates - I've been running this on several machines for years with no issues.
ZFS needs to be kept in sync with the kernel version. If you use Nix on Ubuntu, Ubuntu manages the kernel, so Nix can’t do anything.
ZFS on NixOS is pretty neat though. The exact version pinning means that ZFS and the kernel can be updated together and only once both packages build at a new version.
There’s also an open issue for more extensive checks, so hopefully eventually NixOS CI will be able to check not just that the ZFS and kernel are compatible (header-wise), but actually go through the motion of creating test ZFS dataset and make sure it actually works.
I tend to use stable for everything then overlay the unstable/git master versions of the packages I need.
EDIT: Actually ignore me, of course stable won’t have problems because updates are held back….
I don’t think the CI thing will help. I think what’s happening here is a build time check that prevents any new kernel version with ZFS until the upstream announces support. I don’t think this check will be removed even if there was better CI.
Also, possibly your issues might go away if you used ‘config.boot.zfs.package.latestCompatibleLinuxPackages’ (from NixOS wiki for ZFS). I think that version is only updated once upstream announces support for a newer kernel.
OpenZFS was created as a way to combine development of all the versions of ZFS across Linux, FreeBSD, Illumos, and others. If you're using a (fairly) recent Illumos, you're using OpenZFS.
Disclaimer: Not an OpenZFS developer; just a happy user.
The whole history is pretty convoluted, since I'm fairly certain "OpenZFS" has been used to refer to several completely different source trees, with the Linux, FreeBSD, OS X, NetBSD, and Windows ports of ZFS all being separate forks off of the ZFS code from early OpenSolaris releases.
Despite the licensing drama and general distaste most distros and Linus have/had for ZoL, the Linux port had sort of become the de facto "standard" for ZFS development, with more work being done on that version of the source tree.
I only followed the "drama" from a distance, don't use ZFS or FreeBSD very frequently, but I know there was an internal discussion a while back within FreeBSD where they decided to pull commits directly from the ZoL tree instead of Illumos, because that tree was more active, and otherwise they had to wait on any improvements from Linux to be merged to the upstream Illumos tree.
The elephant in the room was the lack of activity in Illumos and the gradual decline of industry sponsors and development over the years. I don't know what the current state is, and I'm not saying Illumos is "dead", but it's definitely a shadow of its prime back when Joyent and a lot of companies were really pushing and developing on it. Several big names dropped out of Illumos development, including some NAS companies IIRC.
Anyway, OpenZFS 2, AFAIK, is essentially a rebranding of ZFS on Linux, but with CI/CD pipelines testing every commit against FreeBSD as well, basically unifying the work for Linux + FreeBSD.
The other versions (Mac, NetBSD, Windows, others?) still use separate forks AFAIK.
Not sure what the Illumos guys ended up doing in response, either. My memory from browsing their mailing lists was they were pretty surprised by FreeBSD's decision, and didn't seem particularly happy about it, since Illumos/OpenSolaris is the "proper home" of ZFS, but...
This actually happened because of one feature: ZFS encryption. That feature was developed for ZoL first, and had to be ported to other platforms. The developer of the feature for ZoL did not want to write it for Illumos first and port it to Linux later.
The consequence of that is that the Linux port got its first significant feature the others didn't have, and there was no real assistance to bring it into the other platforms quickly. That led to a bit of a political mess where the fallout was the "new" OpenZFS codebase where Linux and FreeBSD are maintained in one tree.
Source: I observed the whole thing while working at the place that wrote ZFS encryption for ZoL.