Native Encryption for ZFS on Linux
github.com
github.com
We originally ran OpenSolaris _just for ZFS_. Then we switched to FreeBSD when OpenSolaris/OpenIndiana dead-ended. We had a production server running PHP and we were sick of PHP5, so we switched to Linux because HHVM on FreeBSD was a joke. PHP 7 came out and obviated HHVM (from a performance perspective), and so our latest PHP deployments are back on FreeBSD.
If your data is valuable to you (and presumably it is and that's why you're looking at ZFS), why are you not on FreeBSD in the first place?
Same story with ASP.NET - tried deploying on Mono way back before .NET Core was even an idea and realized the futility of it and switched a server farm to Windows _because it just wasn't worth it._ PHP on Windows? Same story - not production ready, move to Linux.
(This isn't to disparage ZfsOnLinux, which I think is a great effort and laudable if only for home user purposes. Instead, this is a question for sysadmins on HN who are using ZoL.. .why?)
This isn't just applicable to Linux; we have a strict policy on using only in-kernel filesystems on FreeBSD as well after some poor experiences (_in production_) with aufs-esque ports to FreeBSD. They all look shiny and nice on the outside, and even stress test OK, but when you release them to the masses that's when the shit hits the fan and you realize they're just not designed to the same specifications as the rest of the OS components.
I'm saying it didn't matter why. Just that at the end of the day, the kernel team isn't opting their weight behind this, and that should factor in to the decision.
EDIT: I would also like to explicitly point out the false equivocation to which I am responding. The idea that the Linux Kernel project ignoring the code because it has an incompatible license is, in any way, equivalent to the Linux Kernel Project maintainers choosing to not "throw their weight" behind a piece of code by mainlining it -- essentially that they are rejecting it due to lack of technical merit -- is specious at best and deliberately instils fear, uncertainty and doubt at worst.
If you have ever written a kernel module out of tree then tried to upgrade the kernel version you know this pain.
ZFS was shipped on FreeBSD in 7.0 -- about a decade ago. That's 10 years of development, testing, QA, and use in production.
Canonical just recently began providing kernel modules for ZFS. It appears as if Red Hat may never include support for ZFS and -- unless some changes happen WRT licensing -- it will almost certainly never be part of the Linux kernel.
I've had "ZFS on root" setups on my laptops and workstations on Ubuntu (previously) and Arch Linux (currently) and have several servers using ZFS on FreeBSD. Initial installation and setup of ZFS is a major pain in the ass on Linux, compared to FreeBSD -- especially when dealing with anything more complicated than a single ZFS pool on a single disk.
On funtoo/gentoo/arch basically any distro where the installation is manual zfs is exceptionally simple to set up. This stage took me about 30 seconds to figure out having never used zfs.
On any rolling release you could easily be using the same setup for the next decade or more simply duplicating an existing setup to a new machine periodically.
A small amount of work up front seems like a good trade for never having to do so again.
I never had any issues with this on FreeBSD. Linux, however, was another story. Installation would go smoothly and without incident but I'd run into issues immediately upon first boot. I spent many hours trying to figure out what/where the problem was and it turned out to be race conditions with systemd and filesystems/partitions being mounted. (At the time, I had a few years experience with ZFS and had installed who-knows-how-many servers using ZFS exclusively, yet I had never had any issues with this. Of course, none of those servers were infected by systemd either.)
Now that I'm aware of the issues I can work around them. If you just want a single pool on a single disk and a minimal number of datasets, you'll probably be fine -- you're right, that's a cinch to set up. Like I said, however, don't be surprised if you have problems when trying to do anything more complicated than that.
I'm one of those people. I'm on the Perf and OS team at Netflix. We're very familiar with FreeBSD -- we use it on our CDN. We're also very familiar with Linux -- we use it on our cloud. We're very familiar with a lot of other OSes too (Windows, AIX, Solaris, HP-UX, etc).
So why don't we use BSD everywhere, or Linux everywhere? There's no single reason -- it's many reasons -- and it's serious engineering time to document them all and do the topic justice. I'd like to, but I don't have that time right now.
But I wanted to point this out because I do see it commonly asked, and it sounds like a question that can be answered simply, but isn't. You're asking for engineering work, even to just summarize all the factors. I've written up an explanation of why not Solaris before on HN, since that's easy. But Linux vs FreeBSD is much more work, since they are both compelling options.
I'm saying use whatever filesystem Linux provides when you need to use Linux. And use whatever operating system provides native support for ZFS when you need to use ZFS.
We're using ZFS on FreeBSD to do our data storage, ASP.NET on Windows to run our business logic, and PHP under FreeBSD where PHP is required (previously HHVM on Linux because HHVM was the requirement). I'm saying pick the right OS for the right job.
I'm saying figure out your priority with stability in mind.
ZFS is a must when managing big data. It's nice to have otherwise. Decide which factor is the "must" in your case, and if it's ZFS, don't use Linux.
That is why people risk their data using an untested hack.
1. Is ext4 a decent general purpose file system?
2. Is ext4 a decent alternative or comparable to ZFS?
My answers to those questions would be:
1. It's an "acceptable" general purpose file system but it's really starting to fall behind the competition. When I see comments like how people don't need features like snapshots et al I'm reminded about just how popular Windows Shadow Copy is. Ok that is a service rather than a file system feature but it shows that there is a demand for desktops to have this feature. Checksuming is another example - it's practically free service these days as CPUs have hardware extensions for popular checksum hashes. So the argument is akin to saying "why do we need file system journaling"; sure you don't need them but you're massively grateful when you do have them and your consumer hardware throws it's inevitable hissyfit.
In my opinion Linux really does need to up its game. Ext4 feels almost stuck in time when compares to the likes of XFS which is years older than ext and yet they're working on snapshotting. Then you have really forward thinking file systems like what DragonflyBSD have been working on. They understand the need for better resilience of our data on desktops due to ever expanding storage capacities and HAMMERFS looks extremely exciting as a result.
So to get back to your question. Yes ext4 is acceptable for most people, but is acceptable really good enough? Maybe we should invest more time and energy into XFS?
2. Is ext4 a decent alternative or comparable to ZFS? Simply put, no.
* handles crashes - yes
* detects corruption - no
I'd want all three for a general purpose FS.
Snapshotting is nice but it's complex and expensive, and you can do fine without it. But telling you if your data is safe should be an expected feature for a general purpose FS.
ext4 is a perfectly fine filesystem for most desktop users and a wide range of servers simply by not being ZFS.
If you need CoW and checksumming, there is btrfs too, which is natively supported on Linux and IMO fine if you're only on a single disk or use md or LVM for RAID.
It's quite amusing tbh when people think that a filesystem needs to do everything.
Checksumming and snapshots etc. give substantial game-changing improvements to the whole system. More than a journaling filesystem ever did. They also provide means as to vastly simplify other tasks. Roll-back last system-wide update? Done. Backup consistency? Free and vastly simplified.
I disagree, these are properties that are expected of any decent general purpose filesystem. If you fall into a niche that doesn't, fine, use one of the many alternatives. But this should be the default for all user-facing machines as well as servers.
Journaling is rather useful because especially consumer desktop computers tend to be mistreated and crash a lot. The physical journal in ext4 prevents the worst when stuff gets hairy. Additionally, it does not effectively delete data like ZFS (or rather, make it very very hard to access data) just because a bit was corrupted.
I think a user will value their family pictures more than some random text file suddenly having another character somewhere, they will loose data with a checksumming FS on a normal machine. There is no second drive to pull a good copy of the data from. There is one. And you just killed the last picture of grandma. Congrats.
>Why waste clock cycles on a journal when you might not need it?
Journaling wastes mostly bandwidth and IO, not much CPU and clock cycles.
>Roll-back last system-wide update? Done. Backup consistency? Free and vastly simplified.
That is indeed something you can do with a CoW system but the tools to achieve this are in my experience not fully mature yet.
The same tasks can also be achieved without snapshots at all, NixOS seems to be doing fine.
>Backup consistency? Free and vastly simplified.
Snapshots are not backups. If you pretend they are you will loose all your data.
Snapshots are snapshots. Nothing more and nothing less. Not backups.
>But this should be the default for all user-facing machines as well as servers.
I've not stated they shouldn't be default, rather, that the end user does not necessarily need them.
Though, ZFS is doing a rather poor job on all these tasks, especially disk management on ZFS is a hassle I wouldn't have a typical end user face. I think Bcache FS is a much cleaner and better design than ZFS, which is rather unflexible.
But again, I don't see the immediate need for every end user machine to have a checksumming and CoW filesystem, people have been doing fine for years without them and I doubt there is chaos and fire everywhere as you seem to imply. It could be a little bit better but ext4 is a perfectly fine filesystem for production data.
If you rely on checksumming and snapshots for your data safety, you're in for a bad time. Proper backups, not snapshots, will beat snapshots any day. Checksums make data inaccessible, a rather bad choice for a user machine. Whoops, I guess those family photos are totally unusuable because 1 bit has corrupted.
>Snapshots are not backups. If you pretend they are you will loose all your data.
The good thing about snapshots it that you can use them as a consistent base for your backup: If you naively copy files over from a running system, you'll get inconsistent data, but if you copy it over from a snapshot, it'll work without taking the system down.
If you still need snapshots, there are several methods without needing CoW Filesystems, notably LVM and dattobd (both of which I have utilized in the past)
Also keep in mind that a pure snapshot from a running system will behave like a system that just crashed, a half written file won't complete the write magically because you have snapshot.
If you don't tell postgresql that you're doing snapshots or backups, both will likely corrupt data, doesn't give you anything in that case.
Same for desktop systems. You make a snapshot of the runnign system and suddenly the browser only finds corrupted settings because a write wasn't complete until after the snapshot.
And as far as I know most software acts responsibly with files, either writing out the new version to a temp file or using sqlite. Both of which are atomic-snapshot-safe.
Just because a checksum doesn't match and you don't have redundant copies doesn't mean you throw away data (hint: you don't in practice either). I think a user would rather know about corruption rather than, ehm, not.
Depending on FS you can also disable checksumming, happy?
Of course a user would rather know about corruption, but on the average user system with family pictures, you don't run into problems even after years to my experience. The photos function perfectly fine with minor corruption, which as noted above, does not work on ZFS since on ZFS it's either perfect or repairable or lost.
>I've explicitly said that snapshots HELPS with backups, not that they are backups.
You can do Snapshots on other FS' too, LVM and dattobd allow Snapshots of Blockdevices fairly easily. So I still don't see why that makes ZFS superior to ext4 when you can simply layer the solution underneath.
Also, btw, LVM has RAID1 and RAID6, both of which allow the user to detect data corruption on any filesystem, which IMO, is better than restricting to one filesystem. That way you can make an image of an old FAT16 device and be assured that it won't suddenly corrupt data.
Same goes for mdraid and snapraid, the later of which is a file-based solution and allows recovery of data beyond loosing all parity of an array plus additional drives. ZFS does to my knowledge not work well once you loose more than it tolerates.
Why choose ZFS when I can layer the solution together? Or rather, layer until a FS with a better underlying design comes around.
[1]: https://docs.oracle.com/cd/E26505_01/html/E37384/gbbbc.html
ZFS is targeted towards the enterprise and such is adapted for that use case. I agree that it isn't perfect for home users and I'm not advocating that ZFS be the acceptable solution. I'm just stating that Linux does not have a decent general purpose filesystem.
There is nothing that says that a checksum mismatch results in data loss, that statement is absurd. ZFS has the stance that you should not work on corrupted data as that will result in further corruption when action are based upon it. I tend to agree with that approach, but ZFS will not remove the data it finds corrupted.
You can also layer a more traditional filesystem upon ZFS to leverage redundancy, checksumming and snapshots, something that mdadm does not provide. There are very good reasons for why a layered approach isn't nearly as flexible, this was controversial at the time and debated alot a decade ago.
Most RAID solutions don't even verify the data when read from normal operations. A checksummed approach is a significant improvement, in combination with redundant information you can automatically recover - something that is hard to do in a layered approach. Also, with with a checksum you can verify the correct result. With raid 5 or degraded raid 6 you can at best know that an error has occured but not fix it. Again something that a layered approach has issues with (even if the filesystem had checksumming the RAID controller/software wouldn't know or be able to act upon it).
We're now using btrfs too.
What's your technical beef with ZFS on Linux anyway?
To be honest I do agree with him. I find FreeBSD to be close enough to Linux that any backend engineer or operator worth their salt should be able to swap between the two platforms. Nearly all the same tools that are available on Linux are available on FreeBSD, plus a few tools of it's own. I'm not advocating FreeBSD everywhere though - Linux has it's strengths too which would make it a better platform in other domains. But sometimes it feels like people turn to ZoL without even considering FreeBSD - which is a real pity as FreeBSD + ZFS is a dream to administrate. It's a lot easier than people expect and by not even entertaining the idea of FreeBSD I think they're missing an opportunity.
This is all ideal world scenario stuff though. I do appreciate there are often other factors that complicate the decision making. Sometimes even political rather than technological factors.
ZFS works well enough on either these days, and most of our linux folks are certainly comfortable enough in BSD-land, but our primary server monitoring stack at the time (NewRelic's server/infrastructure agents) doesn't have any BSD support at all and they didn't have interest in adding it, so we went back to the other set of tools we knew (linux) which integrated better with the other tooling we used and didn't have time to replace.
The purists who try to suggest ZFS should be used from BSD whenever possible are probably forgetting "perfect is the enemy of good"
I feel your pain with NewRelic as I've also ran into a few brick walls with them. In fact on Linux as well (not tried NewRelic agents on FreeBSD).
In terms of using FreeBSD, I generally work in small (~6 person) sysadmin teams running hundreds of servers. There's already a lot of software and OSes to maintain and not a hell of a lot of time to do it, if we can cut out some of the complexity (making sure provisioning tools work on FreeBSD etc.) by just running as much stuff on Linux as possible then our jobs and lives become easier.
If you ever have the time, that would be an awesome read. I find the differences between operating systems fascinating, and rarely done justice as people overlook the subtleties for a "Just use <<insert favourite tribe>>" approach.
Zfs can be made to work, shoehorned in your parlance, and then you have the best filesystem on the very best of operating systems.
I too started mid 90s with Linux. Walnut creek cdroms anyone?
I believe Linux merits stands on it's own feet and I don't want to have to write an essay about everything that is awesome about Linux.
- kernel dev speed
- the disruptive gplv2 linux kernel license
- kernel breadth of drivers
- great hardware support
- kernel subsystems available
- container technology in particular lxc
- scalability
- openness
- now - zfs - with encryption support!
- strace
- kvm
- size of community
- size of pool of corporate backed devs
- great great wm's
- great userland tools
- bedrock distributions like Debian
- distributions pushing the envelope like Arch
- The Arch Wiki
What? Manage a git clone, develop until done, submit patches. Btrfs could have done this, but what is the benefit over saying the feature isn't ready? Nobodies forcing you to use or even install it.
The biggest points you could probably win an argument on is:
1) hardware support: but that's less of an issue with servers anyway. Plus ironically I've had more issues installing Debian onto HP Proliants than FreeBSD due Debian's strict policy on non-free drivers.
2) size of community: but then what's the signal to noise ratio in the Linux community of actual helpful, knowledgeable engineers vs that of FreeBSD or any other POSIX-like platform? Linux does have a significantly bigger community but I've found that does result in significantly more misinformation being published too.
A genuine, non-trolling and non-elitist question: have you actually tried any other POSIX-like platform aside Linux? I'm not saying you should need to since you're clearly happy with Linux, but since the points you praise Linux can be applied to a number of platforms outside of Linux as well it does sound like you're making assumptions that those points are exclusive to Linux as they are why you're describing Linux as the best OS out there.
Because the filesystem isn't the only feature of the OS that matters, and people often have more than one requirement. There are technical reasons to choose linux that have nothing to do with ext or xfs, and there are nontechnical reasons also. Just because someone has chosen linux doesn't mean they don't value their data.
Also consider hardware support, mindshare and the flexibility/need to use more of the OS than just the filesystem.
I began with ZFS on Solaris, then Nexenta... then jumped to ZFS on Linux in 2012. FreeBSD was a non-starter because my hardware platform (HPE ProLiant) was poorly supported under FreeBSD. That killed it as an option.
I am seeing long standing stability/availability bugs finally getting fixed within the ZoL project - often times being fixes to OpenZFS upstream, fixing bugs that certainly "felt" like the same thing on those other platforms.
There are zero platforms I entirely trust to survive every conceivable drive failure, I've had production outages using every combination of hardware and OS possible due to bugs both in the OpenZFS stack itself, as well as underlying OS issues on all platforms.
At this point I think it's choose which OS you're most comfortable with, the ZFS bits are getting really close to on par.
And there are many reasons to want to run a linux kernel with ZFS - not just "preference" related.
This is good news, but I'll definitely want to wait a good long while before enabling this in production. Yes, officially zfs isn't good enough to use in production anywhere even without shiny new features, but I reckon zfs as-is is better than some other filesystem.
In theory, having compression and encryption handled by one codebase might open the door to a safe/clearly delineated tradeoffs of enabling both.
I don't know if most of you are old enough to remember the absolute HELL of dealing with tape drives, and UFS to trying to get data recovered or migrated before ZFS. It usually resulted in tears.
> I hope that FreeBSD will be not too far behind. …
Besides, in 99% of "typo instances" it is quickly obvious that a typo was made, everyone understands what it should be, and continues on. If it's a typo that actually results in a real problem then by all means point it out; otherwise it just adds to the noise (such as in this case, where it was merely a transposition of two characters).