ntfs2btrfs: In-place conversion of NTFS filesystem to Btrfs
github.com
github.com
I can safely say it has not presented any problem with me thus far, and I am at the stage of my life where I realize that I don't have the time to fiddle as much with settings. If the distributions are willing to take that maintenance on their shoulders, I'm willing to trust them and deal with the consequences – at least I know I'm not alone.
I second this however I don't use the filesystem to get this functionality. I most often use XFS and have a cronjob that calls an old perl script called "rsnapshot" [1] that makes use of hardlinks to deal with duplicate content and save space. One can create both local and remote snapshots. Similar to your situation I have used this to fix corrupted git repos which I could have done within git itself but rsnapshot was many times easier and I am lazy.
https://kb.synology.com/en-us/DSM/tutorial/What_was_the_RAID...
For scrubbing, plain-old btrfs-scrub is being used.
and they'd be correct
However, the checksum must be chosen at the time of filesystem creation. I don't know of any way to upgrade an existing BtrFS filesystem.
Contrast this to ZFS, which allows the checksum to be modified on a file-by-file basis.
Bitrot still happens with ECC.
Having filesystem level protection is good, but it its like going from 80% to 85% protected. That is because the most critical part remaining unprotected in a traditional RAID/etc system is actually the application to filesystem interface. Posix and linux are largly to blame here because the default IO model should be async interface where the completion only fires when the data is persisted and things like read()/write()/close() should be fully serialized with the persistence layer. Otherwise, even with btrfs the easiest way to lose data is simply write it to disk, close the file, and pull the power plug.
For example if your use case is a file archive (think raw photo or video), then fiesysytem interface does not matter - if the computer crashes soon after copy, you re-copy from original media. But bit flips are very real and can ruin your day.
But as I mentioned these days I'm pretty sure nearly all the storage loss that isn't physical damage is actually software bugs. Just a couple months ago I uploaded a multiple GB file (on my ECC protected workstation) to a major hyperscaler's cloud storage/sharing option. Sent the link to a colleague half way around the globe and they reported an unusual crash. So I asked them to md5sum the file and they got a different result from what I got, so I downloaded the file myself and diffed it against the original and right in the middle there was a ~50k block of garbage. Uploaded it again, and it was fine. Blame it on my browser, or whatever if you will, but the end result was quite disturbing because I'm fairly certain my local storage stack was fine. These days i'm really quite skeptical of "advanced" filesystem/etc. What I want is a dumb one where the number 1 priority is data consistency. I'm not sure that is an accurate reflection of many of them, where winning the storage perf benchmark, or feature wars seems to be a higher priority.
I never found the cause, because I just switched to a completely different system to copy data. I know it was not disk-specific because this was happening on multiple hard drives, nor it was physical damage, as SMART/syslog were silent, and reading the disk again was giving correct data. Memory was fine -- not ECC, but I did run a lot of memtest's on it.
Later on, I found some blog posts which mentioned the similar problem and claim it was result of bad SATA card, or bad cable, or even bad power supply. I remember there was original one made by Jeff Bonwick on his ZFS blog, but I cannot find it anymore. Here is a more modern link instead: https://changelog.complete.org/archives/9769-silent-data-cor...
I now have the homegrown checksumming solution which I use after each major file transfer, and I have not seen any data corruption yet (*known on the wood().
I think the answer is its all covered.
Looking at even SATA 1, if it is implemented _CORRECTLY_ you get the packet CRC protecting the link data from the point its formed to the point the endpoint verifies the result. So like ethernet sometimes it doesn't matter if some random piece of junk in the process doesn't do its own ECC validation, because its covered under a higher level of the stack.
If your adapters are "desktop" grade then I might consider seeking another vendor if you care about data integrity, some vendors are definitely shipping crap, but there are vendors I can assure you will detect link/etc failures.
And as a side, note, I've seen a lot of bad data, a huge percentage of it, were kernel/etc filesystem errors. We added a bunch of out of band extra metrics to track when/where writes were going and our own metadata layers/etc, and it uncovered a whole bunch of software errors.
>>it is designed with a focus on data integrity by protecting the user's data on disk against silent data corruption caused by data degradation, power surges (voltage spikes), bugs in disk firmware, phantom writes (the previous write did not make it to disk), misdirected reads/writes (the disk accesses the wrong block), DMA parity errors between the array and server memory or from the driver (since the checksum validates data inside the array), driver errors (data winds up in the wrong buffer inside the kernel), accidental overwrites (such as swapping to a live file system), etc.
Also tagging is better than writing it down. Like initials+date.
Don't insult people's intelligence then command them to spend hours to prove that you're wrong.
Just start with the assumption that you may be wrong next time.
Also I wouldn't trust BTRFS for the purpose of data archiving, for the fact that ext4 is a proven (and simpler) filesystem, thus it's less likely to became corrupt, and it's more likely to being able to recover data from it if it will become corrupt (or the disk has some bad sectors and that sort of stuff).
On the contrary; I'm using btrfs and not ext4 on NAS (Synology) specifically, because the former does checksumming and bitrot detection and the latter does not.
The backup I was referring to is the offline one. If I need to backup data, something that unfortunately I don't do as often as I should, I need a filesystem that is reliable, that is proven (ext4 is around since more than a decade, and if we count the previous version even more, so in 20 years I'm confident that I would be able to mount an ext4 hard drive that I forgot in the garage in a modern system, with BTRFS, who knows), and that for it they exist a lot of tools in case something goes wrong (there are ton of tools to recover data from damaged ext4 drives, are we sure that with BTRFS is as easy? If I have a filesystem with compression recovering data I don't think is a simple as running photorec...)
Also, the filesystem of a backup drive is not something you can change easily. I still have an old 1Tb drive that I formatted long long time ago in NTFS, and I never changed the filesystem since having to backup all the data to another drive (find it another 1Tb drive), format the drive, and copy back the data will take 1 day. Not that there are things super important on that drive, mostly is stuff I downloaded from the internet years ago, still it's an example why for a backup drive I don't want to have the cutting edge choice that then creates problems in the future.
Ext4 is ubiquitous, so it's my filesystem of choice for all the purpose that have the requirement that the data must be archived for more than 2 years.
Quite a lot of the assumptions of earlier file systems is the hardware either returns correctness, or reports a problem e.g. uncorrectable read error or media error. That's been shown to be untrue even with enterprise class hardware, largely by the ZFS developers, hence why it exists. And also why ZFS has had quite a lot less "bad press" where Btrfs wasn't developed in a kind of skunkworks, it was developed out in the open where quite a lot of early users were using it with ordinary every day hardware.
And as it turns out, we see most hardware by make/model doing mostly the right things, a small number of make/models, making up a significant minority of usage volume, don't do the right things. Hence, Btrfs has always had full checksumming of data and metadata. Both XFS and ext4 were running into the same kinds of problems Btrfs (and ZFS before it) revealed - torn writes, misdirected writes, bit rot, memory bit flips, and even SSD's exhibit prefail behavior by returning either zeros or garbage instead of data (or metadata). XFS and ext4 subsequently added metadata checksums, which further reinforced the understanding that devices sometimes do the wrong thing and also lie about it.
It is true that overwriting filesystems have a better chance of repairing metadata inconsistencies. A big reason why is locality. They have fixed locations on disk for different kinds of metadata, thus a lot of correct assumptions can be made about what should be in that location. Btrfs doesn't have that at all, it has very few fixed locations for metadata (pretty much just the super blocks). Since no assumptions can be made about what's been found in metadata areas, it's harder to fix.
So the strategy is different with Btrfs (and probably ZFS too since it has a fairly nascent fsck even compared to Btrfs's) - cheap and fast replication of data via snapshots and send/receive, which requires no deep traversal of either the source or destination. And equally cheap and fast restore (backwards replication) using the same method. Conversely, conventional backup and restore are meaningfully different when reversing, so you have to test both the backup and restore to really understand if your backup method is reliable. That's going to be your disaster go to rather than trying to fix them. Fixing is almost certainly going to take much longer than restoring. If you don't have current backups, at least Btrfs now has various rescue mount options to make the file system more tolerant of broken file systems, but as a consequence you also have to mount read-only. Pretty good chance you can still get your data out, even if it's inconvenient to have to wipe the file system and create a new one. It'll still be faster than mucking with repair.
Also, Btrfs since kernel 5.3 has both read time and write time tree checkers, that verify certain trees for consistency, not just blindly accepting checksums. Various problems are exposed and stopped before they can cause worse problems, and even helps find memory bitflips and btrfs bugs. Btrfs doesn't just complain about hardware related issues, it'll rat itself out if it's to blame for the problem - which at this point isn't happening any more often than ext4 or XFS in very large deployments (millions of instances).
I didn't talk only about corruption of the filesystem itself (I don't know if it's more or less likely with BTRFS, someone says that BTRFS is more likely to become corrupt with power failures, I don't know if it's true), but also from hardware failures. In case of a disk with damaged sectors (I know that we should have 3 backups with one offsite, but you always have that one disk with important data on it that it's a year that you are promising to backup next day till it breaks) I think that a filesystem with a simpler structure will lead to an higher probability of recovering the data, while I think that with BTRFS, or any filesystem that is COW, uses compression, volumes, etc that is more difficult, because files are not stored as plain blocks on the disk, but have a more complex structure that must need to be decoded.
Also BTRFS is kind of a new filesystem, that has two disadvantages, there are not all the tools that were developed over the years for ext4, and also BTRFS driver is continuing evolving. Why I can be pretty confident that if I format an hard disk today with an ext4 filesystem in 20 years I will find a driver for a modern Linux (or whatever OS will replace it in 20 years) to mount it, can we have the same assurance with BTRFS? I don't know.
So for the purpose of making backups and archiving data, I think that I will stick with ext4 for a while. While on my laptop, and systems that I use, I use BTRFS without any problems.
No. If the drive honors flush/FUA, Btrfs is less likely to corrupt data or metadata than overwriting file systems because the interruption won't result in incomplete overwrites. So this would hold true for any copy-on-write vs overwriting file system (and probably also log based file systems). The trouble is if the drive is transiently lying about flush/FUA success, and then there's an ill timed crash. There's the chance the super blocks written point to trees that don't exist yet because the write order hasn't been honored due to flush/FUA being ignored. There are backup trees, so it might be possible to work around this defect with the `rescue=usebackuproot` mount option, but sometimes the defect is so bad that you get all kinds of write reordering such that Btrfs only finds trees with the wrong generation, and it fails to mount. Often it's still possible to get your data out with the offline scrape tool, `btrfs restore`. But it's a difficult problem to deal with. In theory it's similar on ZFS but I know nothing about its on-disk format so maybe its metadata has some locality in which case certain assumptions could be made to allow it to better work around such a drive firmware defect? I'm not sure. On a power fail, it is possible Btrfs loses the most recently written data if the writes that were in-progress and thus not yet fully committed to stable media. How much data really depends on the application doing the writes.
>In case of a disk with damaged sectors
Btrfs by default keeps two copies of metadata and it automatically deals with this problem, while also self-healing when such problems are encountered.
>a filesystem with a simpler structure will lead to an higher probability of recovering the data, while I think that with BTRFS, or any filesystem that is COW, uses compression, volumes, etc that is more difficult, because files are not stored as plain blocks on the disk, but have a more complex structure that must need to be decoded.
The ondisk format is fairly simple and extendible. Metadata isn't subject to compression. In the case of bad sectors with compressed (user) data, you'll certainly lose more data than if it weren't compressed. There's an expected trade off here, it's not really a Btrfs issue but just the way all compression algorithms work. You get some small corruption and it'll have a bigger effect.
>So for the purpose of making backups and archiving data, I think that I will stick with ext4 for a while.
I used to hedge my bets by having multiple copies of data on different file systems (including ZFS) but haven't done that in years. I've seen too many cases of (hardware induced) data corruption being replicated into backups and archives without any warning it was happening until it was too late - and only corrupt copies remained.
BTRFS sounded cool with all it's new features, but the reality is that ext4+LVM does absolutely everything I need and it's never given me any issues.
I'm sure BTRFS is much more robust these days, but I'm still gun shy!
A week later, the power failed in the office, and my filesystem got corrupted. In the middle of repairing, the power dipped again, and my entire fs was unrecoverable after that. I managed to get the data off by booting off a separate drive and running a command to extract everything, but it would never mount no matter what I did.
I've never had an issue with ext4, xfs, or zfs no matter how much I've abused them over the past 10+ years, but if losing power twice can wipe out my filesystem then no thanks, I'm out.
(Plus: non-recursive snapshots? No thanks.)
EXT4 has never once failed me, and I personally battle tested it by working while losing power probably a total of 200 times. I probably should have bought a UPS come to think of it.
These were single disk systems, no raid at all.
Around the timeframe you mentioned, I lost a BTRFS filesystem when I filled it up. Probably could have recovered it if I had known more, but oh well. I definitely feel the gun-shyness!
However, I'd want to add that at a previous job, I had a ext4 fs go belly-up in a similar way. One day, just died without warning. Maybe could have recovered it, but like others have mentioned we'd have no guarantees about the data.
Moral of the story is, of course, always have backups :)
On the other hand, those days might be behind it. I haven't kept track.
The upstream maintainer of btrfs-progs also maintains btrfs maintenance https://github.com/kdave/btrfsmaintenance which is a set of scripts for doing various tasks, one of which is a periodic filtered balance focusing on data block groups. I don't run it myself, because well (a) I haven't run into such a bug in years myself (b) I want to run into such a bug so I can report it and get it fixed. The maintenance script basically papers over any remaining issues, for those folks who are more interested in avoiding issues than bug reporting (a completely reasonable position).
There's been a metric ton of work on ENOSPC bugs over the years, but a large pile were set free with the ticketed ENOSPC system quite a few years ago now (circa 2015? 2016?)
Does any other FS on Linux provide those?
But... any of the other stuff? Checksums, deduplication, compression, etc. Citations, please.
LVM VDO wasn't really a thing until a couple of years ago, and I've never actually heard anyone recommend using it. Definitely not within its "traditional domain".
Perhaps you mean on other operating systems... when this discussion is all about a Linux filesystem and whether other Linux filesystems provided these features, including the comment you replied to?
Yes. The Linux LVM logical volume manager for example can do those things.
Traditional domain meaning this kind of data management traditionally came under the purview of LVMs. Not that an implementation has had particular features for a length of time, but that you might (also) look there rather than the filesystem for those features.
As fair as I know the only solution for deduplication in LVM is VDO and that was created after ZFS and isn't technically part of LVM (it's a kernel module that sits between the filesystem and LVM).
Same is true for compression except in the case of compression that's been a file system feature since the days of NTFS on NT 3.5 (released 1995) -- I can't recall when UNIX file systems first saw compression but I'd wager it was before ZFS, which we've already acknowledge predates VDO.
I also thing it makes sense that the above should be part of the file system domain because it's an integral part to how the file system organizes the data (much like how we don't argue that the journal isn't part of the file system domain).
As for checksums, they have existed in all parts of the stack, from the application level (for example rsync), to the volume level (RAID controller cards) to the file system. I wouldn't say that's something that has ever had a "traditional domain".
The things I'd say ZFS and its ilk do that isn't part of its traditional domain is RAIDing (ie management of the physical devices) and cache control. Both make a lot of sense being managed by the file system in modern file systems but I can also sympathize with those who think that's a step too far. In the case of RAIDing, you can still run ZFS on LVM or a hardware controller if you wish -- I wouldn't suggest you do so because it gives you additional complications for zero benefit but the option is still there if you want it. But with regards to cache, the only way to opt out of ZFS managing it's own cache is not to use ZFS.
Anyway it wasn't really a statement about what was there first or not, but rather that because something may not exist for a filesystem does not mean it does not exist, i.e., look at other layers for such functionality.
https://www.redhat.com/en/blog/look-vdo-new-linux-compressio...
One can also try convert a desktop installation into a server installation (not sure if it’s possible)
That said, I'm thinking about leaving a BTRFS partition unmounted and mounting it only to perform backups, taking advantage of the snapshotting features.
Device replace
--------------
>Device replace and device delete insist on being able to read or reconstruct all data. If any read fails due to an IO error, the delete/replace operation is aborted and the administrator must remove or replace the damaged data before trying again.
Device replace isn't something where "mostly ok" is a good enough status.
I mean I'm all for choice on Linux but when it comes to file systems I'd rather have fewer choices but those choices be absolutely rock solid. Btrfs might have gotten to that stage now however it's eaten enough peoples data (including mine) over the years that I can't help wondering why people bothered to persist with using it when a better option was available.
Try to fill your root-filesystem with dd, then remove the file and sync, reboot and enjoy a non booting OS ;) It's like they don't test it at all.
But then they make changes that add an insane amount of complexity, and suddenly you're running into random errors and googling all the time to try to find the magical fixes to all the problems you didn't have before.
I eventually gave up (/ got sick of doing restores) and just copied the data into a fresh btrfs volume. That worked "great" up until I realized (a) I had to turn off CoW for a bunch of things I wanted to snapshot, (b) you can't actually defrag in practice because it unlinks shared extents and (c) btrfs on a multi-drive array has a failure mode that will leave your root filesystem readonly; which is just a footgun that shouldn't exist in a production-facing filesystem. - I should add that these were not particularly huge filesytems: the ext3 conversion fiasco was ~64G, and my servers were like ~200G and ~100G respectively. I also was doing "raid1"/"raid10" style setups, and not exercising the supposedly broken raid5/raid6 code in any way.
I think I probably lost three or four filesystems which were supposed to be "redundant" before I gave up and switched to ZFS. Between (a) & (b) above btrfs just has very few advantages compared to ZFS. Really the only thing going for it was being available in mainline kernel builds. (Which, frankly, I don't consider that to be an advantage the way the GPL zealots on the LKML seem to think it is.)
ZFS doesn't have defrag, and BtrFS does.
There was a paper recently on purposefully introducing fragmentation, and the approach could drastically reduce performance on any filesystem that was tested.
This can be fixed in BtrFS. I don't see how to recover from this on ZFS, apart from a massive resilver.
https://www.usenix.org/system/files/login/articles/login_sum...
"Btrfs, ext4, F2FS, XFS, ZFS all age - up to 22x on HDD, up to 2x on SSD."
https://pdfs.semanticscholar.org/b743/7111bf04a803878ebacbc2...
The paper also mentions ZFS:
I don't know, I'm just hoping for a filesystem that can get these features right to come along...
That was pretty funny, and I agree a thousand times over. When I was younger (read: had shallower pockets) I was willing to spend time on these hacks to avoid the need for intermediate storage. Now that I'm wiser, crankier, and surrounded by cheap storage: I would rather just have Bezos send me a pair of drives in <24h to chuck in a sled. They can populate while I'm occupied and/or sleeping.
My time spent troubleshooting this crap when it inevitably explodes is just not worth the price of a couple of drives; and if I still manage to cock everything up at least the latter approach leaves me with one or more backup copies. If everything goes according to plan well hey the usable storage on my NAS just went up ;-). I feel bad for the people that will inevitably run this command on the only copy of their data. (Though I would hope the userland tool ships w/ plenty of warnings to the contrary.)
Buying an extra disk for just the conversion is wasteful, and then you need space to keep it stashed forever when you never use it. Not at all sustainable, I'd rather leave the hardware on the market for people who _actually_ need it.
Now that's really interesting.
We did the rsync, started the new system, it seemed to be working okay, but then we started seeing some weird errors. After some investigation, it looked like the rsync didn't work right. We were tired, it was getting late, so we decided to put one of the original mirrors in the new system since we knew it worked.
Started up the new system with the old mirror, it ran for a while, then started acted weird too. At that point we only had 1 mirror left, were beat, and decided to pack the old and new system up and bring it all back to the office (my co-founder's house!) and figure out what was going on. We couldn't afford to lose the last mirror.
After making another mirror in the old system, we started testing the new system. It seemed to work fine with 1 disk in either bay (it had 2). But when we put them in together and started doing I/O from A to B, it corrupted drive A. We weren't even writing to drive A!
For the next test, I put both drives on 1 IDE controller instead of each on its own controller. (Motherboards had 2 IDE controllers, each supported 2 drives). That worked fine.
It turns out there was a defect on the MB and if both IDE ports were active, it got confused and sent data to the wrong drive. We needed the CPU upgrade so ended up running both drives on 1 IDE port and it worked fine until we replaced it a year later.
But we learned a valuable lesson: never ever use your production data when doing any kind of upgrade. Make copies, trash them, but don't use the originals. I think that lesson applies to the idea of doing an inplace conversion from NTFS to Btrfs, even if it says it keeps a backup. Do yourself a favor and copy the whole drive first, then mess around with the copy.
Other than whatever that instability was, I can say that the performance was exceptional and would use that setup again, with more investigation into causes of the instability.
Main usecase is storage spaces w/o server or workstation.
[1] ...given how broken it seemingly is, see features that are half baked like raid-5. But I am a ZFS snob so don't mind me, my fs of choice has it's own issues.
Regardless, in-place conversion is specifically a feature of btrfs due to how its designed. Since it doesn't require a lot of fixed metadata, you can convert a fs in-place by throwing the btrfs metadata into unallocated space and just point to the same blocks as the original fs. I think it even copies the original fs's metadata too, so you can mount the filesystem as either the original or btrfs for a while.
I also was misled by btrfs-check which happily scanned my un-mountable filesystem with no complaints, and learned on the irc that btrfs-check is "an odd duck," in that it doesn't actually check all the things the documentation says it does.
This experience, and the fact that simple things like the "can only mount a degraded filesystem rw one time" bug remain after years of complaints, simply because the devs adopt an arrogant "you're doing it wrong if you don't fix your filesystem the first time you mount it" (despite the tools giving you no indication that's required) attitude, have convinced me to never touch btrfs again.
So keep parroting "it's stable!" all you want, my experience has shown btrfs is "stable" until you have a problem.
"Stable" here means "unchanging". It doesn't mean bug-free. My personal experience is, you'll generally encounter less bugs on a rolling release distro (like arch) or distros with frequent updates (like fedora). The upside of stable distros (like debian) is that new bugs (or other breaking changes) won't be introduced during the distro's release lifetime.
> So keep parroting "it's stable!" all you want, my experience has shown btrfs is "stable" until you have a problem. I've been running it on multiple production machines for years now, as well as my home machine. Facebook has been using it in production for I think over a decade now, and it's used by Google and Synology on some of their products.
I'm not saying it doesn't have problems (I've certainly faced a few), but it is tiresome reading the same cracks against it because they set it up without reading the docs. You never see the same thing against someone running ZFS without regular scrubs or in RAIDZ1.
A well-designed, general-audience technology product doesn't require one to be initiated into the mysteries before using it. The phone in my pocket is literally a million times more complicated than my first computer and I haven't read a bit of documentation. It works fine.
If btrfs wants to be something people use, its promoters need to stop blaming users for bad outcomes and start making it so that the default setup is what people need to get results at least as good as the competition. I have never read an extfs manual, but when I've had problems nobody has ever blamed bad outcomes on me not reading an extfs manual.
Its window to it was when setting up ZFS included lots of hand-waving. That window now has closed. ZFS is sable, does not eat data, does not have a cult of wizards spelling "RTFM" and is can be installed in major distributions using easy to follow procedure. In a year or two I expect that procedure to be fully automated, to a point where one could do a root on ZFS.
I do use the Ubuntu live image pretty regularly when I need to import zpools in a preboot environment and it works great. In general it's not my favorite distro - but I'm happy to see they're doing some of the leg work to bring ZFS to a wider audience.
[1]: https://openzfs.github.io/openzfs-docs/Getting%20Started/Ubu...
Ubuntu has been able to install directly to root-on-ZFS automatically since 20.04. I don't think any other major distros are as aggressive about supporting ZFS due to the licensing problem, but the software is already there.
ZFS doesn't have these kinds of hidden gotchas, and that's the key difference. Yeah ok somebody's being dumb if they never scrub and find out they have uncorrectable bad data come from two drives on a raidz1. That's exactly the advertised limitation of raidz1: it can survive a single complete drive failure, and can't repair data that has been corrupted on two (or more) drives at once.
If you are in the scenario, as the GP was, that you have a two-disk mirror and regular scrubs have assured that one of the disks has only good data, and the other dies, ZFS won't corrupt the data on its own. If you try replacing the bad drive with another bad drive, eventually the bad drive will fail or produce so many errors that ZFS stops trying to use it, and you'll know. The pool will continue on with the good drive and tell you about it. Then you buy another replacement and hope that one is good. No surprises.
Why is ZFS requiring scrubs and understanding the limitations of it's RAID implementations okay, but btrfs requiring scrubs and understanding the limitations of its RAID implementations "hidden gotchas"?
> If you are in the scenario, as the GP was, that you have a two-disk mirror and regular scrubs have assured that one of the disks has only good data, and the other dies, ZFS won't corrupt the data on its own. Honestly, I don't know enough about GP's situation to really comment on what happened there. It could have been btrfs or perhaps they were using hardware RAID and the controller screwed up. ZFS is definitely very good in that regard and I want to be clear that I'm not saying ZFS is bad or that btrfs is better then it; I've been using ZFS much longer then I have btrfs, back before ZoL was a thing.
It read more like btrfs corrupting data despite good scrubbing practice; hosing the file system on the good drive instead of letting it remain good, for instance. If that's a misreading, that is where my position came from.
Regular scrubs and understanding the limitations of redundancy models is good on both systems, yes.
My own anecdotal evidence though: btrfs really does snag itself into surprising and disastrous situations at alarming frequency. Between being unable to reshape a pool (eg, removing a disk when plenty of free space exists) and not being safe with unclean shutdowns, it's hard to ever trust it. It even went a few good years where it seemed to be abandoned, but I guess since 2018 or so it's been picked up again.
>Between being unable to reshape a pool (eg, removing a disk when plenty of free space exists) and not being safe with unclean shutdowns, it's hard to ever trust it. It even went a few good years where it seemed to be abandoned, but I guess since 2018 or so it's been picked up again.
FYI, btrfs does support reshaping pool with the btrfs device commands.
I've found it trivially easy to get btrfs stuck in a state where the commands to do so refuse to function.
My favorite is #8885.
https://github.com/openzfs/zfs/issues/8885
that's (a) not really a bug in ZFS, and (b) "fails to boot sometimes" is pretty different from btrfs shitting the bed and corrupting its pool. There was one of those recently with ZFS iirc (and specifically only ZFS-on-Linux) but they are fairly rare and notable when they occur!
(As a general statement, ZoL is less mature than ZFS-on-FreeBSD and likely (perhaps) to continue to be so given the licensing issues. I've also run into some problems where I can't send a dataset from FreeBSD to a ZoL pool (but rsync works fine). But again, generally bugs that actually lead to data loss are exceedingly rare.)
That's strange since FreeBSD actually moved to ZoL a while ago
Even if you grant it to be a simple bug, it's not exactly a grave design issue with ZFS.
From what I saw, it happened if the disk initialization took longer and ZoL looked for its pools before all disks were found. It hints at improper dependencies in ZoL startup.
But if you're using it in some other setup, then that means you went out of your way to try a more complicated filesystem. I would think it's reasonable to do at least a quick scan of the btrfs wiki or your distro's documentation before continuing with that, the same way I'd expect someone would do the same for ZFS.
I would argue that the defaults should not be dangerous. If a filesystem is released in to the world as stable and ready for use, it's not absurd to expect that running mkfs.<whatever> /dev/<disk> will get you something that's not going to eat your data. It might not be optimized, but it shouldn't be dangerous.
Dangerous defaults are loaded footguns and should be treated as bugs.
If there is no safe option, there should not be a default and the user should be required to explicitly make a choice. At that point you can blame them for not reading the documentation.
There are some arguable footguns in the btrfs-tools workflows for repairing a damaged filesystem, but that's exactly the fragile situation where asking the user to RTFM before making more changes is perfectly reasonable.
The ones being discussed in the parent posts in this thread. I am not myself particularly familiar with btrfs internals, having only used it once, but I have heard of there being issues along those lines.
> There are some arguable footguns in the btrfs-tools workflows for repairing a damaged filesystem, but that's exactly the fragile situation where asking the user to RTFM before making more changes is perfectly reasonable.
My point is that in the cases where RTFM is a requirement there should not be a default behavior. Doing the dangerous thing should always require an explicit request and not be something that one can autocomplete their way to. If there is a "doing X is only safe when you also fill in Y parameter and put Z in mode W" then it either shouldn't let me do X without those other things at all or should require a "--yes-really-do-the-stupid-thing" type flag.
The thing mentioned upthread where it's possible to mount a damaged filesystem RW, but only once. If that's true, then attempting to do so through a normal command someone who knows Unixy systems might just perform without thinking should scream bloody murder to make sure the user knows what's going on, and then should require some explicit confirmation of intent to move forward with the dangerous and/or irreversable operation.
(I've been a happy user of btrfs for several years now https://news.ycombinator.com/item?id=21446761)
What I'm trying to convey is most of the time when I help someone out with btrfs issues it always turns out they haven't been running regular scrubs or balances and their distro didn't setup anything so they didn't know. Or they don't understand how much diskspace they have because their distro's docs still have them running du instead of "btrfs filesystem usage".
For specific distros, arch/arch-derivatives had autodefrag as a default mount option which caused some thrashing a while back and IIRC Debian still doesn't support subvolumes properly, as well as not installing btrfs-maintenance.
As far as I can tell, that is true. But given the number of years it spent broken, the amount of time I spent recovering broken arrays, and the (unknown but >0) number of my files that it ate, you'll forgive me for not being enthusiastic now.
Who fiddles with ext4 or xfs after formatting it?
You thought that'd be a swipe but that's how the developers pronounce it