BTRFS is frequently not considered stable enough for production usage.
ZFS has dozens of useful features besides RAID. Transparent compression, instant atomic snapshots, incremental snapshot sync, instant cloning of file systems, etc etc.
Yes, different ZFS implementations are mostly compatible in my experience, and they should become totally compatible as everyone moves to OpenZFS. FreeBSD 13 and Linux currently have ZFS feature parity I believe.
synology NAS use btrfs by default. What is not stable?
I'm not sure this is still true, especially after Facebook deployed it on millions of servers.
For example, detecting corruption would be more important than never having any, so that you can get IO errors and immediately kill the server before it serves bad data.
https://www.phoronix.com/scan.php?page=news_item&px=Btrfs-Wa...
Others have touched on the main points, I just wanted to stress that an important distinction between ZFS and hardware RAID and linux software RAID (by which I assume you mean MD) is that the latter two present themselves as block devices. One has to put a file system on top to make use of them.
In contrast, ZFS does away with this traditional split, and provides a filesystem as well as support for a virtual block device. By unifying the full stack from the filesystem down to the actual devices, it can be smarter and more resilient.
The first few minutes of this[1] presentation does a good job of explaining why ZFS was built this way and how it improves on the traditional RAID solutions.
Some highlights: hierarchical checksumming, CoW snapshots, deduplication, more efficient rebuilds, extremely configurable, tiered storage, various caching strategies, etc.
- You generally want to avoid hardware raid, if the card dies you'll likely need to source a compatible replacement vs. grabbing another SATA/was expander and reconstructing the array.
- zfs handles the stack all the way from drives to filesystem, allowing them to work together (i.e filesystem usage info can better dictate what gets moved around tiered storage, or better raid recovery.
Apparently in the enterprise server versions it's fairly decent, but the desktop versions are pretty trash re: performance.
Hardware RAID is actually older then ZFS style software RAID. ZFS was specifically designed fix the issues with hardware RAID.
The problem with Hardware RAID is that is has no ideas what going on on top of it, and even worse, its a mostly a bunch of closed-source fireware from a vendor. And they cost money.
You can find lots of terrible story about those.
ZFS is open-source and battle tested.
> linux software RAID
Not sure what you are referring too.
> BTRFS does RAID too.
BTRFS is basically copied many of the features done in ZFS. BTRFS has a history of being far less stable. ZFS is far more battle tested. They say its stable now, but they had said that many times. It eat my data twice so I have not followed the project anymore. A file system in my opinion gets exactly 1 chance with me.
They each have some features the other doesn't but broadly speaking they are similar technology.
The new bcacheFS is also coming up and adding some interesting features.
> Why would people choose ZFS in 2021 if both Oracle and Open Source users have 2 competing ZFS?
Not sure what that has do with anything. Oracle is an evil company, they tried take all these great open source technologies away from people and the community thought against it. Most of the ZFS team left after the merger.
The Open-Source version is arguable better, and has far more of the original designers working on it. The two code bases have diverged a lot since then.
At the end of the day ZFS is incredibly battle tested, works incredibly well at what it does. And had a incredible reputation of stability basically since it came out. They question in my opinion is why not ZFS, then why ZFS.
Did you mean "it ate my data" to apply to ZFS? Or did you mean BTRFS?
I never fell for the BTRFS meme but many friends of mine did, and many of them ended up with a corrupted filesystem (and lost data).
1. Do you trust Firmware...i don't, i can tell you storys about freaking out san's...never had that with solaris or freebsd and zfs.
2. Why having a additional abstraction layer, HW Raid caching vs FS-Caching, no transparency for error correction, not smart raid rebuild etc.
the list can go on and on, but HW-Raid is a thing of the past (exceptions are specialized san's etc)
Hardware RAID controllers predate ZFS by a long time.. ZFS is a much more modern design and because it integrates the whole storage layer it can offer all the features it does, which a RAID controller hiding behind a disk interface cannot do.
When ZFS came out many people (me included) considered that the end of relevance for hardware RAID controllers. I used to use hardware RAID pre-ZFS but have never again after switching to ZFS when Solaris 10 first included it.
RAID does not feature the data protection offered by a copy on write filesystem, and OpenZFS is the most stable and portable option.
Btrfs is not stable in all configurations. Mdadm and others don't do checksumming, scrubs, health checks, etc as well as ZFS (or at all). That's not even touching on built in encryption, compression, snapshots, boot environments and more.
That's just the worst of all worlds: Usually proprietary and you get the extreme aversion to improvement (or any change really) of hardware vendors.
This ZFS changes is going to come, and it may end up being complex to implement for users... but it's happening. At the risk of being hyperbolic: Something like would never be possible with a HW raid system unless it had explicitly been designed for it from the start.
Also: ZFS does much more than any hardware RAID ever did.
With zfs, it has checksumming so it knows which data copy is the correct one.
As drives get bigger, the odds of bit decay, while being low on a per-bit level, become great enough for the drive in total, and is a concern.
File systems have had checksumming, but if there is no exposure into the raid layer, the file system can't use this checksum to recover from bit decay errors since it can't control which drive processes a request.
This is why zfs operates at both layers.
zfs vs btrfs is harder, and i'd guess its the zfs layered read and write caching systems enticing people in.
Today's many-TB hard drives use mind-boggling engineering to shrink data fluctuations down almost to the size of individual atoms (not quite there yet, but making progress (!)), while higher I/O speeds push the onboard processors' ECC mechanisms almost beyond breaking point.
This basically means that the likelihood of HDDs returning uncaught bit errors has gone from a maybe-once-in-a-lifetime event (with very large magnetic flux sizes on disk in the 80s) to something that individual power users using disks with great-looking SMART output should generally expect to see every few months or years+.
My rule of thumb is that any storage device over maybe 250GB-500GB in size needs to be redundant. If you can shove that much data into one place, the odds that something significant will be hiding somewhere in that data easily go beyond 100% once you go past that size range, IMO.
Other than that one extreme double-failure scenario being worked out, BTRFS has proven remarkably stable for a while now. A decade ago that wasn't quite as absolutely bulletproof, but today the situation is much different. Personally, it feels to me like there is a persistent & vocal small group of people who seemingly either have some agenda that makes them not wish to consider BTRFS, or they are unwilling to review & reconsider how things might have changed in the last decade. Not to belabor the point but it's quite frustrating, and it feels a bit odd that BTRFS is such a persistent target of slander & assault. Few other file systems seem to face anywhere near as much criticism, never so out of hand/casually, and honestly, in the end, it just seems like there's some continent of ZFS folks with some strange need to make themselves feel better by putting others down.
One big sign of trust: Fedora 35 Cloud looks likely to switch to BTRFS as default[3], following Fedora 33 desktop lat year making the move. A number of big names use BTRFS, including Facebook. I have yet to see any hyperscalers interested in ZFS.
I'm excited to see ZFS start to get some competent expandability. Expanding ZFS used to be a nightmare. I'll continue running BTRFS for now, but I'm excited to see file systems flourish. Things I wouldn't do? Hardware RAID. Controllers are persnickety weird devices, each with their own invisible sets of constraints & specific firmware issues. If at all possible, I'd prefer the kernel figure out how to make effective use out of multiple disks. BTRFS, and now it seems ZFS perhaps too, do a magical job of making that easy, effective, & fast, in a safe way.
Edit: the current widely-adopted write hole fix is to use RAID1 or RAID1c3 or RAID1c4 (3 copy RAID1, 4 copy RAID1) for meta-data, RAID5/6 for data.
[1] https://btrfs.wiki.kernel.org/index.php/Status
[2] https://btrfs.wiki.kernel.org/index.php/RAID56
[3] https://www.phoronix.com/scan.php?page=news_item&px=Fedora-C...
E.g. Sailfish OS is perhaps the only mobile OS I know that uses / used BTRFS in production (and they adopted it nearly 6-7 years ago!). And some of its users have had issues with BTRFS in the earlier versions - https://together.jolla.com/questions/scope:all/sort:activity... ... in fact, I too remember that once or twice, we had to manually run the btrfs balancer before doing an OS update. For Sailfish OS on Tablet Jolla even experimented with LVM and ext4, and perhaps even considered dropping BTRFS. (I don't know what it uses for newer versions of Sailfish OS now - I think it allows the user to choose between BTRFS or LVM / EXT4).
Most users consider a file system (be it ZFS or BTRFS) to be a really low-level system software with which they only wish to interact transparently (even I got anxious when I had to run btrfs balancer on Sailfish OS the first time worrying what would happen if there was not enough free space to do the operation and hoping I wouldn't lose my data). Even on older systems, everybody frustrated over the need to run a defragmenter.
Perhaps because of improper expectations or configurations, some of the early adopters of BTRFS got burnt with it after possibly even losing their precious data. It's hard to forget that kind of experience and thus perhaps the "continuing hate" you see for BTRFS - a PR issue that BTRFS' proponents needs to fix.
(It's interesting to see the progress BTRFS has made. Thanks to your post, I may consider it for future Linux installations over EXT4. Except for the hands-on tinkering it required once or twice, I remember it as being rock-solid on my Sailfish mobile.)
https://documentation.suse.com/sles/15-SP1/html/SLES-all/cha...
(Details: Corrupted filesystem, happened twice, ~2019 IIRC, on a single disk system so not even touching the RAID code, first time couldn't repair, second some didn't try. Hasn't happened again and wasn't just a checksum error so I doubt that hardware is at fault but could be wrong.)
I also keep a clean copy of an installed debian install snapshot on my NAS so i can just send it to new machines rather than run through the whole setup. works great.
Oh you are talking about ZFS, aren't you. ;)
Playing the "who came up with it first" game doesn't particularly interest me. Again, it feels like a spirit of competition when to me, that seems like a bad spirit: we should be cooperative & boosting each other. We're both open source, we're both trying to make civilization possible & to share greatnesses.
[1] http://0pointer.net/blog/revisiting-how-we-put-together-linu...
Is that why you sound like a sour apple?
No just two Gnu/Linux distributions:
openSUSE/SLES and Fedora
And SUSE (SLES) tells you (strong recommendation) to use btrfs just for the OS (for data XFS, just like RHEL) and just in a mirror configuration.
Look i work since ~forever with SLES, and if you operate it in mirror mode (just os), don't touch it and just make snapshots i never had problems. But i HAD complete data-loss with btrf many times when i did for example a defrag and re-compress, those are native btrfs-tools, and that is not acceptable to me. The Filesystems is THE place in a OS where errors like that are not acceptable (to me).
User of freebsd myself but that is BS, Netflix uses freebsd and UFS2 on the openconnect device.
https://openconnect.netflix.com/en/
You watch youtube about netflix they state clearly that they use ufs.
I "think" it's that one
https://www.youtube.com/watch?v=veQwkG0WdN8
Microservice-backend is linux/aws
We use UFS for content because we rely on zero-copy async sendfile for our high performance video serving data path. When using ZFS, sendfile is not async, and because ZFS uses ARC rather than the page cache, it requires a copy from ARC to network with ZFS, so its yet ready for our high performance workload.
We don't use RAID at all.
Thanks for clarification, so can i assume that bectl/adm is in full swing?
Meanwhile I've been running ZFS for close to the same time and have never lost anything.
I get it's an anecdotal view point but that's a very hard reputation to rebuild for BTRFS.
If the btrfs failed for you, you’ll remember and likely won’t try again. I had that experience with a now defunct hard drive maker. After the second replacement one failed I vowed never again.
They probably should rename it.
Look here is the problem. The BTRFS people declared multiple time that it is stable. For a few years it was repeatably 'BTRFS' is stable now if you use it in such an such a away.
But then it destroyed a drive.
Then it a few years later and now it was actually stable.
But then it destroyed data again.
I know what they are saying about themselves but unfortunately the project has simply lost credibility with a lot of people. File system should be developed to be stable first and then slowly add features never being unstable. They shouldn't still regularly destroy peoples data after years of years of development.
While at the same time ZFS has been basically stable since early days, many developers at Sun switched their root drives to ZFS before it was even officially released. It never had years and years of routine instability.
> Few other file systems seem to face anywhere near as much criticism,
Where few file system claimed they were stable for years while routinely losing data. The file system has one job first and foremost, don't destroy users data.
This is not a conspiracy, this is simply the reality that tons of people have lost data because BTRFS repeatably claimed stability when it wasn't.
Its now another couple years later, likely this wouldn't happen again. But that after all these years they still haven't managed to get RAID 5/6 fully working doesn't exactly scream confidence. So me and I assume many others simply have lost confidence in the project and the approach they take to the development.
Kirk McKusick creator of UFS/2
So once they are burnt, they wont come back.
When I did a fresh install of Fedora 33 on my primary workstation 3 months ago, I had exactly the same rationale for sticking with the default BTRFS selection. "I'm sure it has come quite a long way since I last tried it, I would like to have some of those features, I know there are some large production installations now, and the fact that it's the default in Fedora is a sign of confidence from the community."
After 2 months of use, I ended up with a corrupt filesystem in the middle of my work day and could not find any way to recover from it other than to do a full reinstall and restore my files from a backup. This was on a single NVMe drive in a system running a few small VM's, a browser, a chat client, and a few terminals. Thankfully I only lost a day's worth of work, but that's the last time I install BTRFS on any of my personal systems.
> [...] it feels a bit odd that BTRFS is such a persistent target of slander & assault.
My experience is obviously anecdotal (as are all individual experiences), so I won't be surprised if you dismiss my comment like you did nullwarp's comment. But "slander & assault" just seems like a weird way to dismiss all critics of BTRFS at once, as if everyone is out to get BTRFS. Filesystems have a thankless job. Do it right, and most users will never even think about it. But lose a user's data once, and you've likely lost that user forever.
> A number of big names use BTRFS, including Facebook. I have yet to see any hyperscalers interested in ZFS.
ZFS is excellent for large arrays of spinning disks, but if you're using a bunch of fast SSD's, performance really sucks. There is a lot of lock contention contributing to that which isn't noticeable on slower devices. I can't speak to FB's environment, but if they're managing a large number of SSD's like most hyperscalers, then ZFS would probably get ruled out based on performance comparisons.