ZFS Root Filesystem on AWS
scotte.org
scotte.org
The first main issue was the arc_shrink_shift default being poor for machines with a large ARC. Our machines have Arc at several 100GB, so the default arc_shrink_shift was flushing several GBs to disk at a time. This was causing our machines to become unresponsive for several seconds at a time pretty frequently.
The other main issue we encountered was when we tried to delete lots of data at a time. We aren't sure why, but when we tried to delete a lot of data (~200GB from each machine which each contain several TB of data), our databases become unresponsive for an hour.
Other then these issues, ZFS has worked incredibly well. The builtin compression has saved us lots of $$$. It's just the unknown unknowns that have been getting us.
The licensing is incredibly unfortunate, though. (I don't care about the reasoning for the license, it's just bad that it isn't GPL-compatible so that it could be compatible with the most prolific kernel in the world.)
Anyway, back to BTRFS-vs-ZFS. It seems abundantly clear that a filesystem is (no longer) a thing where you can just "throw an early idea out there" and hope that others will pick up the slack and fix all the bugs. There's just too much design (not code) that goes into these things that it's not just about code any more.
My (small) bet right now as to the "next gen" FS on Linux is on bcachefs[1, 2]. It sounds much sounder from a design perspective than BFS, plus it's built on the already-proven bcache, etc. etc. (Read the page for details.)
[1] https://www.patreon.com/bcachefs [1] https://bcache.evilpiepirate.org/Bcachefs/
This means Linux copyright owners could sue ZoL binary distributers, but Oracle could not.
However no one is shipping ZoL binaries, only the source code. The code itself is 100% conflict free.
There used to be an issue where users hitting their quota couldn't delete files since for some reason deleting a file meant creating a file somehow. The trick was to find some reasonably large file and `echo 1 > large_file` which truncates the file and frees up enough space that you can begin removing files. Maybe this kind of trick could help you guys.
That said, it's inadvisable to run a database on a btree file system like ZFS or btrfs if you're keeping an eye on the write performance. cf Postgres 9.0 High Performance by Gregory Smith (https://www.amazon.com/PostgreSQL-High-Performance-Gregory-S...)
and
https://blog.pgaddict.com/posts/postgresql-performance-on-ex...
Our writes are actually heavily CPU bound because of how we architected the system[0]. We recently made some changes that dramatically improved our write throughput, so AFAICT, we aren't going to need to focus much on write performance in the near future.
[0] http://blog.heapanalytics.com/running-10-million-postgresql-...
1. Reflink copies. Basically this is a file level snapshot. The metadata is unique to the new file, but it shares (initially) extents with the original.
2. Seed device. The volume is set read-only, mounted, add a 2nd device, remount rw, delete the seed. This causes the seed to be replicated to the 2nd device, but with a new volume UUID. Use case might be to do an installation minimally configured so as to be generic, and then use it as a seed for quickly creating multiple unique instances.
Another use case: don't delete the 1st volume. Each VM has two devices: the read only seed (shared), a read write 2nd device (unique). The rw device (sprout) internally references the read-only seed, so you need only mount by the UUID of the sprout.
Seed-sprout is something like an overlay, or a volume-wide snapshot.
I guess it's possible that some type of disk command timing could cause unexpected lockups or slowdowns that you wouldn't get with a system that doesn't try to control the hardware to quite the same extent as ZFS, but my (cursory) understanding is that it's rare/hardware specific.
My personal take is that running ZFS on hardware that lies is no worse than running EXT4 on it. YMMV as I'm not a storage expert.
Might as well treat your zpool like it's on real hardware and configure raidz accordingly. The cloud does have real, problematic hardware behind it, and it's important to remember that.
[edit] especially if you can configure your block devices such that you know they're sitting on different physical hardware at the cloud data center, you will gain that benefit of ZFS
Or in the case if EBS - NFS and a network stack. EBS is "Emulated Block Storage" despite what AWS may like you to think.
https://github.com/zfsonlinux/zfs/wiki/Ubuntu-16.04-Root-on-...
It however advises to use blockid, not device-id, for mounts.
Any idea if this doesn't apply to AMIs?
And maybe people using ZFS want proper volume/subvolume management as good as support for traditional partitions or LVM volumes? If so, it will probably take a while longer to land.
Can you? Sure.
Is it smart to possibly introduce errors into the data path in an otherwise fully checksummed path? Not so much.
If it's imperfect, ZFS will occasionally calculate a checksum wrong and write this to all drives in the array. At some point in the future, like when you read the file, the checksum will fail on all drives and the whole file will be marked corrupt. This gets annoying fast.
Your memory is not error-free, but it might be close enough.
For an analogy, consider a world where car engines have a non-trivial chance to instantly explode when involved in a crash. Then someone comes out with an engine that doesn't explode. People say "you should wear seatbelts if you use this non-exploding engine," but your car has no seatbelts; clearly you are still safer with the non-exploding engine, but all of the sudden, seatbelts are more likely to save your life than before.
A non-checksumming FS would not be vulnerable to this particular issue. On the other hand, would undetected corruption through bit rot be a worse problem? Almost certainly.
And considering how vanishingly unlikely such a scenario is, I do agree with your sentiment.
That said, I'm not sure if my understanding of the issue is complete and would welcome an explanation of the failure scenarios that [very] occasional RAM bit-flips expose ZFS to.
Let's say that you have a 99.9% chance of the scrub running correctly on a big pool with non-ECC memory (0.1% chance of a bit-flip during the scrub). Any single scrub is extraordinarily likely to succeed, but if you run a scrub every day then over the course of a year then your chance of your pool surviving falls to 0.999^365 = 69.4%.
Pick your favorite numbers here, 0.1% failure chance per scrub is probably way high. With five nines your yearly survival rate is 99.6%. But do remember that soft errors are fairly rare in modern servers mostly because they use ECC RAM, you can't look at data with ECC and assume you'll get comparable results by using non-ECC RAM.
In general, if you scrub infrequently you are probably going to be OK. (but then why are you using ZFS instead of LVM?) If you live at high altitude, however - let's say in Denver - you are also facing significantly increased soft error rates. The extra atmosphere at sea level does make a strong difference in shielding, something around 5x reduction in strike events.
On the plus side - the SRAM and some parts of the processor do use ECC internally, which is good because fault rates increase with reduced feature size and increased number of transistors. The CPU is potentially the most sensitive part of the system per unit area, so it's very important to protect against errors there.
And on the other hand - disk corruption or failure probably outweigh those kinds of concerns in practice. But it's not like it's expensive to get a system with ECC. An Avoton C2550 runs like $250. So why take the risk anyway? Your data's worth an extra $100 in hardware.
Heck, you can run ECC RAM on the Athlon 5350 and the Asus AM1M-A motherboard. Boom, ECC mobo/CPU combo for under $100. It's just a little thin on SATA channels. It's a shame there's no "server" version of this board with dual NICs, IPMI, and an extra SATA controller tossed on there.
Since ZFS caches aggressively in RAM it's conceivable that faulty RAM may write corrupt data to disk.
You shouldn't need to in a systemd based distro: journald logs to /run until /var comes up. (and then flushes across) See https://www.freedesktop.org/software/systemd/man/journald.co...
I'm working on a Packer builder type that will support this use case as we speak. The existing AWS `chroot` builder would likely be sufficient, but requires running from within AWS.
So... very exciting, looking forward to hearing more about when it "hits the streets"!
You can probably hit-up Ed via Github or Server Fault if you want to talk to him about it directly.
But it's at a point where it safely stores your data correctly. Perhaps some init scripts fail on boot to import your pool/etc. but the data is there.
We do run it production, but we also have in-house tooling built around it.
ZFS-on-Linux devs say it's ready for production[1].
Lawrence Livermore laboratory stores petabytes of data using ZoL[2].
If we're sharing anecdotes, ZoL has served me fantastically for several years.
[1] https://clusterhq.com/2014/09/11/state-zfs-on-linux/ [2] http://computation.llnl.gov/newsroom/livermores-zfs-linux-po...
1) Seen users complaining about data loss on issues on github. 2) Had the init script fail on upgrade and had to fix it by hand when upgrading Ubuntu. Probably a one time issue.
Need a bit more reliability from a file system.
https://github.com/zfsonlinux/zfs/issues/5535
We're strongly considering using something else until this gets addressed. The problem is, we don't know what, because every other CoW implementation also has issues.
* dm-thinp: Slow, wastes disk space
* OverlayFS: No SELinux support
* aufs: Not in mainline or CentOS kernel; rename(2) not implemented correctly; slow writes
If that were me, I'd see how quickly it was fixed before strongly considering something else.
In our opinion, when we made the switch, it was much more important to trust the integrity of the data, than any possible kernel panic.
I have seen a few servers with ext4/mdraid over the last five years have serious corruption but have had to reset a ZoL server maybe twice.
I transitioned an md RAID1 from spinning disks to SSDs last week. After I removed the last spinning disk, one of the SSDs started returning garbage.
1/3 reads are returning garbage and ext4 freaks out, of course. It's too late and the array is shot. I restore from backup.
This would have been a non-event with ZFS. I've got a few production ZoL arrays running and the only problems I've had have been around memory consumption and responsiveness under load. Data integrity has been perfect.
i'm perfectly willing to believe there may be some rare situations where zfs on linux will cause you a problem. but i bet they're rare enough it'll have saved you a few times before it bites you.
The setup process was a bit painful given some interesting delays when using some HW storage controllers that caused udev to not make some HDD devices available under /dev before the ZFS scripts kicked in and we have been bitten a couple times by changes (or bugs) in the boot scripts, however the gains provided by ZFS in terms of data integrity, backup, and virtual machine provisioning workflow were definitely worth it.