Early Linux filesystem reliability
minnie.tuhs.org
minnie.tuhs.org
Back in ye olden days, yanking the power cord basically guaranteed a corrupt Ext2 FS and a visit from fsck. However, in the vast majority of cases, fsck would actually do the job. You had to peruse lost+found and recover the lost souls out of it, but that was usually the extent of it. I did 'fsck -y' many, many times - in most cases with good results.
XFS on PC was worse. It tended to be stupendously fast when dealing with lots of I/O to/from very large files (video editing, running VMs), but a power failure on PC was a lot worse than with Ext2. There was a good chance you would lose the whole volume.
On Irix / SGI hardware it was a very different story. Those suckers were quite reliable. Heavy as hell, too.
---
One habit from back then that's very hard to shake off is to run sync before reboot, often as "sync; reboot" - you know, just in case something gets stuck on the way down and you have to hit the reset button. A more extreme example would be to manually stop all services except sshd, then do "sync; reboot".
It's completely unjustified today, and yet I still do "sync; reboot" occasionally. It's baked into the muscle memory of my fingers after sleepless nights caused by losing stuff due a flaky driver that froze the system on reboot.
There's more about it in this old usenet thread: https://groups.google.com/forum/#!topic/alt.folklore.compute...
Mass-storage back then was often done with tape decks, which were a lot cheaper than hard disks, and the triple-sync trick was a common firmware trick that tape-drive vendors used to get a cheap 'rewind command' out there for their users, without requiring platform-specific bins for the job.
And if I recall, this was a built-in of the tape-drive itself, which could be used on multiple different systems, and so the triple-sync being used to rewind wasn't Risc/OS specific; I remember using the same tape-deck later on Linux and Irix machines, same ol' sync command worked every time.
I can still hear the whirr in my mind ..
Filesystems can be remounted read-only via the serial console as well, with a break-u. Useful even when userland is otherwise inaccessible such as in the event of a fork bomb.
I'm not terribly familiar with KVM but the VM console tools probably have some way to generate a break signal. Check the docs for how to "send a break."
Oh, and don't forget to enable this feature via sysctl -- per the above linux doc.
In Ubuntu/Kubuntu there was once a key combination to forcefully unmount all file systems. THey seemed to have discontinued this? No?
- Alt+SysRq+u remounts all filesystems readonly.
Looks like some of the Magic SysRq Keys are disabled by default now in new Ubuntus.
Before you think I'm too old-fashioned, it's at least the shutdown command and not "init 6"
Especially given how often I have to hard power-off machines because they just freeze, or because shutdown -h now doesn't work, or because they keyboard input doesn't work ...
Eventually the Ext family caught up with it.
Somehow, I remember filesystems to be more reliable by them, and the time needed for the fsck being the biggest issue. Then we had a Red Hat person demonstrate a forced reboot in a computer running ReiserFS in my university, and I think in the next month no computing student was suing Ext2 anymore.
Right, that was always a big problem.
Can you or someone say what was it in the design that allowed XFS to have that performance with large files? Was this mostly extent based allocation or something else?
One time on my personal machine, someone talk bombed me once, and I somehow had to not only recover the superblock from about the 8th backup location (yeah, the first 7 or so failed), but I also had to manually recover some files based on the inode data. But it worked!
Almost all development teams tried so hard to build smart software doing the right thing, recovering from everything, tried to fix so many things. They spent insane amounts of work trying to be smart in error situations. And all of that was useless or harmful - it either didn't work, or it it made things worse.
My team operated under simple principles. Crash early, crash hard, log well, trust the operator. But you know, Ops at that place loved our applications. It took time to get them going, but it was easy to understand what to do whenever they crashed. It was easy to understand why it stopped. It was easy to understand when to file a bug.
Failing early and visible is always good, especially before a program have been taken into production. When things have stabilized for a while, recoverable errors can be ignored, and just logged.
It's also important to take the opportunity to improve the error handling and logging first, before fixing the actual problem. Errors are hard to fake "right" so getting a good reproducible error is an opportunity.
I went from BTRFS to XFS because I got some really bad corner case of performance in BTRFS. My work consists of mostly backend development using Ruby web stack and PostgreSQL/Redis, and often my laptop would freeze completely and only get back after a dozen of seconds. This being in a SSD was unacceptable*. So I decided to go back to XFS and for some time everything went well, no more random freezes.
In this weekend I updated my system by running pacman, as usual. Chrome seemed to consume all memory and the system freeze during the update. Ok, not bad I thought, I rebooted and tried to run pacman again, hopping that only some files was corrupted, however sufficient files were corrupted that I needed to reinstall all packages in the system. Another freeze during the update and my system was essentially death after reboot. I tried to recover using chroot however multiple files were broken beyond repair.
So I decided to go back to EXT4, and reading this article does make me more confident that this shouldn't happen again.
When (if?) Arch supports ZFS in main repos I will probably test it. That is, unless bcachefs comes first.
Not really: zfs is available in aur as a dkms package. Pacman automatically rebuilds dkms modules when it updates the kernel. All you'd need to do is modify mkinitcpio.conf to include "zfs" in the HOOKS at the appropriate stage.
So: more complicated that just using ext4, but not "maintain my own kernel" level of difficulty.
We did have one incident recently with an Ubuntu 14.04 system, which had RAID-1 across 3 drives. Lost one physical drive, and thereby lost the entire btrfs filesystem. Running the btrfs fsck wasn't able to fix it. I likely should have run the latest btrfs-tools to try to fix it, instead of letting the default version that came with 14.04 version try.
Still, we're not planning on switching anytime soon. Been using btrfs send/receive for snapshot backups, which is awesome.
SUSE / openSUSE has had btrfs as the default filesystem for a few years (and we have a bunch of tools built around it adding features like boot-to-snapshot and auto-snapshot of upgrades). Personally I have had issues with it, but I've also messed around with btrfs subvolumes quite a lot (developing container runtime storage drivers) so it might be self-inflicted.
(1) I've only seen this on one machine, so maybe it's a quirk of that machine's workload. But it sure does suck when it's in the middle of a workday.
I recently started using `snapper` to create snapshots on a schedule on them. I enabled quota support in btrfs so I could see how much space snapshots were using.
I noticed that filesystem wide latency tended to spike when removing snapshots (several minutes of all fs access stalling).
Balancing with quotas enabled is even worse: my systems were hung for multiple days, until I forcibly restarted them and disabled quotas. Then the fs hangs were much smaller (a few seconds) and not to noticeable. Balancing finished in something on the order of an hour.
While I had quotas enabled, I was constantly having btrfs tell me the data was bad and needed rescanning (rescanning quotas would also induce fs wide latency).
The thing is, ZFS has snapshot space usage info, and doesn't have awful latency (it also doesn't have a "balance" operation, but I'm not sure how relevant that is).
Given my experience with both btrfs & ZFS, I'll likely consider using ZFS as my rootfs in the future.
I have three machines running openSUSE, and the one I've seen it on is both the least powerful one (Lenovo IdeaPad 100, some Atom chip, 2GB RAM) and the only one with a non-SSD drive.
The main thing that I think is stopping widespread adoption is that none of the developers seems to want to come forth and say "Here's a stable, rock solid version. Go nuts.". There's always a caveat or a disclaimer. Use at your own risk and all that.
When even the developers are kinda jittery about it, it isn't exactly reassuring.
BTRFS for a 'simple' use case is fine.
Push redundancy to another layer for the time being. EG: MDADM or hardware raid.
The main issue for whether it's the default filesystem is whether the distro has the resources to support it for their users. Mostly this is in terms of documentation, and understanding what sort of backports to support. Really the only distro doing that work is Suse and maybe Oracle. I don't expect the more conservative distros to support it for some time until they feel they can depend on upstream's backporting alone.
Had a LOT of unplanned downtime due to various issues with older kernel versions, but 4.10+ has been solid so far. You definitely need operational tooling (monitoring, maintenance like balance) and a good understanding of the internals (what happens when you run our of metadata space etc.).
Happy to answer questions!
On a related note: Never ever use the ext4 to btrfs conversion tool! It's horribly broken and causes issues weeks later.
Care to give some details about this and other failures? Part of what makes a FS reputation is not just people telling "it works" but also stories about how the thing crashed and how they recovered from it. IOW, it always works, until it doesn't, and then it still "works" because I can dig myself out of the hole this or that way.
Inspired by the way you can convert a Debian VM to Arch Linux on Digital Ocean, I happen to have been toying with it recently to auto-convert a blank Debian 8.x VM from ext4 to btrfs. Looks like things are fine, but only because the kernel is <4.x and the VM has very little data on it since it's blank.
WARNING: This is a toy. Do not use for production.
Someone on #btrfs said that the filesystem layout is a lot different when using the conversion tool and all of the regression testing happens with regular filesystems, not converted ones.
We reinstalled all machines from scratch. Never happened again.
Once something gets a reputation for having issues that reputation tends to stick pretty much forever. I'm thinking maybe older versions of btrfs were problematic and the FUD has never gone away.
I remember having a issue with BTRFS two years ago when go a unexpected powerdown (the SAI don't help us as was a short-circuit after the SAI!) where a HyperV VM with BTRFS had the FS corrupted, but we managed to recover the data. Also we have a physical server to run Jenkins and GitLab that it's using FakeRaid + BTRFS + btrbk to schedule backups using snapshots and btrfs send, that are stored as compressed files on a network folder (and this folder are backuped to magnetic tapes).
I don't noticed any slowdown when I launched a manual rebalance, but we are operating at low scale so we not have a lot of I/O. Our real bottleneck are the databases that are on Windows servers.
Time to market matters. Probably explains how codebases with low reputed quality still seems to win in the real world : MySQL vs Postgres ? MongoDB vs RethinkDB ?
The biggest Linux defect was not ext{2,3,4} but LVM and MD, which would throw away write barriers until something like kernel 2.6.31 (which was especially painful on Ubuntu LTS). Many distros used LVM by default, and many servers used mdraid somewhere in the stack. I saw many corrupt Linux systems in the 2000s through the first part of this decade.. it was especially egregious for DBs and hypervisors with file based disk images.
This is insane, I refuse to believe it. Even a junior EE knows how to design a PCB so that the RESET signal is asserted on all ICs as soon as the voltage drops below a safe operating level.
My thought was that the power supply should guarantee adequate power, or none at all. Not something in between. Also, no a rapidly alternating states of adequate power / no power.
Short term solution: rebuild that system and get it a UPS.
Longer term solution: get a much better box.
That is very hard to do, actually. Voltage does fluctuate to a certain degree, because physics. And you don't want to cut power to a system just due to a small power dip. And you can't distinguish a small power dip from 10 small power dips over some time ago. And your flapping detection only works if the power drops are close enough to each other.
Proper control theory is amazingly hard. As my prof on that said - no one fully understands PID controllers, but some blokes have a really lucky thumb.
This can't possibly happen even if the power supply is crazy bad, because the reset logic on the main board will halt everything before the DRAM starts malfunctioning.
When I read it, I got the impression he was saying the DRAM was going crazy while a DMA transfer to the hard-drive was still going on. That doesn't require the CPU to be functional when the DRAM is corrupt, only the DMA controller. I can't personally say if that makes it any more likely though.
Not sure what the hard drive would do with a truncated ATA command though.
ext3 predates SATA, briefly. While SATA was commonly used with ext3 soon after ext3's release, a lot of the hardware Ted Ts'o is probably referring to would have been ATA-4/UDMA (or SCSI).
> while the CPU is still executing instructions
The problem Ts'o described doesn't involve the CPU:
DRAM tends to go insane and starts returning
garbage long before the DMA engine and the hard
drive stops functioning.
A bus-mastering ATA (or SCSI) controller - possibly on the motherboard, possibly an expansion card, regardless almost certainly PCI - may be copying data from RAM directly as Direct Memory Access.> the reset logic on the main board will halt everything
In theory, there is no difference between theory and practice. In practice, there is, especially when cheap, poorly-designed hardware is involved.
"There’s nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem." -Matthew Ahrens (Cofounder of ZFS at Sun Microsystems and current ZFS developer at Delphix)
I think comes from an assumption that if you use ZFS you care more than users of other filesystems about getting the same data you have put onto disk out of the disk later when reading it back.
If that's important it's good advice to use ECC RAM. The point that's often lost is that if you cared about correct data, then having ECC RAM would have been a good idea regardless of using ZFS or not, and you're not worse off with ZFS than you would have been with another filesystem without ECC RAM.
Another common myth is that ZFS requires large amounts of RAM. It doesn't do that either, but the assumption is that if you use ZFS you're building some large multi-user system which benefits from a large cache. If that's not your use-case then that's not true, you could run ZFS on your single disk laptop if you wanted to.
Zfs is no more likely to be damaged that anything else.
This whole confusion stems from zfs devs suggestion that ecc ram be used and people who don't know any better thinking oh that must be some special requirement for zfs.
These folks then spread this misinformation all over the Internet where its passed on by people who know nothing about the topic.
Pro tip: don't talk about things on the Internet that you only know about 3rd or 37th hand. If you know nothing about the topic you aren't improving the store of human knowledge by passing on noise.
If it were "nonsense" then one could flip the argument and say "ECC memory has no beneficial effect what-so-ever" - which be both know is untrue. However I can see why people do say "nonsense" given the ECC-myth seems to only be discussed in relation to ZFS and the probability data errors due to non-ECC memory is small.
For what it's worth, I did originally include a comment about how the risk being the same as on any other file systems - but then deleted it fearing I might have overlooked something on another fs I'm less familiar with.
> Pro tip: don't talk about things on the Internet that you only know about 3rd or 37th hand. If you know nothing about the topic you aren't improving the store of human knowledge by passing on noise.
I don't think that's fair. The problem with 1st hand experience is that it's often just based on anecdote. Which can often be worse than 3rd hand advice. And in the case of failures: often some products are too widely used to be worth constantly badgering the development team for help; which means you end up having to rely on the advice of others. The problem is really more that some people are terrible at researching so take 3rd hand advice without bothering to fact checking it.
In any case, I don't know if your comment was aimed at me or not, but I do have nearly 10 years of experience (wow that's gone quick!) running ZFS across a variety of systems, some running ECC memory, others not. I've also been a keen study of the Sun/Oracle docs. So I do consider myself reasonably well informed from both credible sources and personal experience. Though I'm not arrogant enough to assume I'm an expert either - communities like HN can be deeply humbling places.
> If it were "nonsense" then one could flip the argument and say "ECC memory has no beneficial effect what-so-ever" - which be both know is untrue.
The original argument is like an implication ("no ECC => don't use ZFS"), which he refutes by saying (correctly), that file systems are generally affected the same way by memory corruption. Your argument seems to invert the original implication and drawing conclusions from that (fallacy of the converse, I believe). ∎
I feel between this post and you're previous one that you are basically just agreeing with me via the process of nitpicking the language I used.
Furthermore backups and raid may be related but they are pretty orthogonal. Having a reasonable strategy to ensure performance integrity and reliability isn't obviated by having a good backup strategy.
If you have to restore from backup aren't you going to lose some amount of data even if it's only the last hour -> day.
https://www.linux.com/news/learn/intro-to-linux/how-facebook...