ZFS is mysteriously eating my CPU
brendangregg.com
brendangregg.com
It seems the whole thing took place in 2017 and was fixed then (first by not calling the reclaim if used ARC memory was zero, then the root cause also fixed by using a PRNG instead of the default CSPRNG in 0.7).
No code.
https://ocw.mit.edu/courses/aeronautics-and-astronautics/16-...
as is tradition
On Ubuntu, some kernel packages install also the zfs modules package. I think (but not 100% sure) that if they're on that O/S, with those packages installed, then the zfs module will be loaded on startup.
Is ZFS removable there without recompiling?
https://papers.freebsd.org/2019/fosdem/looney-netflix_and_fr...
# kldunload zfs.ko
should do the trick. You would of course have to change the install setup to use UFS rather than a zroot.I suspect it’s probably loaded by default regardless of OS drive file system, though.
The bottom of the article notes that his book Systems Performance 2nd Edition is 45% off until 9/13.
Edit: looks like the comments about "didn't I see this a few years ago?" have been deleted.
Also, any idea how something like safari online impacts those numbers?
Thanks again for the book, it's an amazing read that I recommend to anyone that is interested in the area.
How much pirated PDFs hurt these books I'm not sure, maybe a little, maybe a lot. With my BPF book, a "rough cut" was published on Safari and then that did the rounds as the PDF (and still does) even though it was unfinished and buggy. That's what annoyed me the most about it, was that people were reading a broken version and may be thinking it's the final version.
Thanks for getting and recommending it!
:-)
Had a colleague who asked for SSH access to production machines to debug an issue. Ops team asked what he wanted to do, guy just wanted to look at which env vars were set. Ops team told him how to do his job - log the configuration - instead of give him access, because they had a mandate to ensure five or seven nines uptime. Can't risk it.
I had a lot more respect for the ops people there than I had for my fellow SWE's.
There are always bugs that will happen only on the prod machines. Sure, they are rare, but they exist.
> just wanted to look at which env vars were set. Ops team told him how to do his job - log the configuration
Well, that's not risk free neither. That needs a code change and you risk exposing secrets, or logging more than you should. (Though I agree with the general idea of you shouldn't be SSHing if you can do it another way)
"do your job" is a really rude way to refer to debugging by logging instead of debugging interactively.
(Another interpretation is that the author previously held Ops in higher regard than SWE and that this event did not change that.)
Dogma results in religious wars
I’d still rather not have to debug arbitrary mutations to the env or file system in a production container though.
It seems to me that shelling into a prod container/VM is discouraged not because you might cause it to fail, but because you might produce undefined behavior while claiming it is healthy (more like Byzantine faults).
For example if you unset a single env var by mistake, then 1/N of your requests will potentially fail. Debugging this kind of issue is a nightmare.
Not to mention that developers can often run arbitrary SQL from a prod shell when the app is backed by a DB.
Or far worse, not fail.
That's a change, inherently much more risky than the "cat /.../env" the colleague wanted to do.
Also the change might well cause the problem to go away, and now you know nothing instead.
The principle is good, but it sounds like it has taken on a life on its own. The ops guy could also have executed the cat command on the spot with the right privileges. Sure, it's gatekeeping, but so is the four eyes principle, and a little of that gatekeeping can be necessary to keep those nines rolling.
Also like others noted here , If one guy sshing into one machine can screw your seven nine uptimes.
Then you never had a seven nine uptime (you’re company is probably just gambling on that uptime metric until some component fails)
But yes, CLI can be dangerous, even for observability. E.g., people running strace(1) on production apps and causing outages due to the strace overhead (I wrote a prior post about that). You need to understand the risks and overheads of all tools. It's why I have a "pull no punches" policy when writing eBPF tool man pages [0]: If the overhead can be bad, it should say so clearly.
[0] https://github.com/iovisor/bcc/blob/master/CONTRIBUTING-SCRI...
So you can log in and debug, but once you're done, the VM is replaced with a clean one.
The only way to really be sure of that is extensive testing at multiple levels, and ideally some form of verification like a TLA+ model.
It might be good for some specific purposes but I'd reckon the majority of people using it do not fall into those. Using it "just because" is probably more trouble than worth.
> ZFS really wasn't in use, ever! But at the same time, it was eating over 30% of CPU capacity! Whaaat??
Interesting bug, sounds like the best option is to not even have the module loaded then. (Not blaming the developers here, necessarily, but yeah, it's a bug)
Regardless, ZFS has configurations equivalent to most "traditional" RAID options anyway (raidz1 is their RAID5) but with the added benefits of all the other great stuff ZFS gives you, like checksums, compression, deduplication, snapshots, etc.
Not using RAID5 or RAID6 modes on Btrfs is the first piece of advice people considering Btrfs get.
[0]: https://www.phoronix.com/scan.php?page=news_item&px=Btrfs-Wa...
On the other hand a lot of problems caused by btrfs is unrecoverable and a mess. Ext4 is stable but lacks the feature set that makes customers go to zfs instead (and that is fine, software should do one thing well and right).
Just that RAID in general and ZFS put my files in a blackbox, and that I didn't really need a significant boost in network reads. I settled for plain old ext4 and periodic rsync mirror files to another disk (I don't mind some interim data loss + writes are fast). Use SFTP or sshfs for accessing the drive.
As someone who has a couple times needed to rip 100tb of ZFS disks out of something, put them into a different machine with a different architecture or OS in some random order, and access everything without having lost anything, it's hard for me to overstate how great it is that such a thing is even possible.
Which ZFS commands?
There’s a potential problem, in that you can’t import a poll on an OS version that lacks the feature flags that are enabled on the pool. The way to solve that is to choose a common subset when creating the pool.
CPU endianness switches work fine, IME - I have pools I originated on SPARC Solaris that work fine on x86, I just got a patch landed for an edge case in one recent feature not interoperating properly between endiannesses but other than that it's worked fine for me.
What kind of issue did you see?
Thats great news, hopefully that rare potential problem got addressed in the last year or so.
Steps to recreate: compile linux (maybe FreeBSD as well, that is a little more fuzzy) for LE. Boot new kernel and setup a new pool. Recompile for BE. Boot new kernel and enjoy bootloop...
lol, actually things are starting to come more into focus... page size played a role in the whole thing. As I said, edge case :) No potential for data loss though.
Sounds like a fun experiment to run down. Maybe I'll go look when I'm done the current nightmare I'm tinkering with.
She swallowed the spider to catch the fly... I'm presently writing something to crack 8051 firmware xnor'd by a 64 byte key, I don't remember how - but the previously described ZFS edge case somehow got me to this point.
I cut a patch to fix sparc64 building with the new zstd feature. I then discovered the new zstd feature was broken for endian portability because lol bitfields.
It ended up re-importing OK when I brought it back to the OpenIndiana box. Whew! I then did an export, and Ubuntu was then able to import the pool.
So hopefully the pool was exported, but the worst case scenario isn't very bad.
That is not going to change between different machines as far as I know, I might be wrong about this. I made some bad choices (in hindsight) when making this pool for my uses, but it's still going strong many years later. [1]
This current pool (which is now old, 50k+ power on hours on these disk) has survived a motherboard dying randomly and one or two SSD failures with the OS on.
Always just installed a OS on a new SSD and it's been picked up just fine.
[1] I went with RAID-Z2 for 6x3TB WD Reds where I probably should have made mirrored pools or something like it to gain more space? It's been a while since I looked at it. Can't really expand this pool or add more storage without replacing each disk one by one with something bigger.
I could make another pool with new disks but I'd lose another 20% to parity.
Caveat: things may have changed since then - it was at least six years ago.
They have not seen a huge amount of reads/writes though, if we don't count the weekly scrub and weekly usage by just me.
This isn't a mirror, not in the sense that most would understand. No performance boost, no protection from bit rot, unnecessarily wasteful in every potential metric. If you wanted to intentionally design a system to propagate errors and render backups useless - you would start with this kind of setup. It makes sense to avoid all the RAID related problems associated with hardware solutions... but I'm drawing a blank on rational reasons to do what you've described. Super weird boot manager + physical space constraints?
If the "live" disk fails, the other disk will eventually replace it.
> render backups useless
Like I mentioned, data loss isn't a concern. This is mainly hosting code repositories and media files.
What's my use-case? My laptop, that has a paltry 256GB SSD, keeps running out of disk space, and I often find myself plonking files that aren't super critical on the network drive.
I knew there was a reason I was getting a "can't send mail more than 500 miles" vibe... Are we talking small form factor spinning rust, or thumb drives? Because if you thought that ZFS was a black box - look into flash wear leveling algorithms, I'm hard pressed to think of any storage media more prone to spontaneous bitflips. Unless you are checksuming before and after every movement of data - your files are going to silently get corrupted in ways you won't notice, until something breaks.
Anyway, I've had way more instances of bitflips than drive failures - and I've had to deal with that kind of data corruption escaping detection and making it into backup archives. That is why I obnoxiously promote ZFS to the degree I do... it totally eliminates that risk automatically. No joke, this thread prompted me to run a manual pool scrub a month in advance for an 8TB archive pool. Since it is archival, it gets few reads - which means it gets few automatic checksum verifications outside of the scheduled bi-monthly scrub. Well the scrub is only half way through, but it already shows that at some point in the last month a 128KB block on a single 3.5" disk got silently corrupted. Data corruption nipped in the bud, thanks to ZFS <insert whatever Sun's jingle was>.
These are old WD Passport spinning hard-disks that were never designed or optimized to run as NAS drives. I had them lying around doing nothing. This is a salvage operation for the old drive.
All that talk of "homelab use-case", "mirrored storage", "RAID", "rsync"... obviously what is under discussion is how ZFS is a poor fit for the tmpfs tier garbage data use-case, dunno how I missed it.