Btrfs has been deprecated in RHEL
access.redhat.com
access.redhat.com
For RHEL you are stuck on one kernel for an entire release. Every fix has to be backported from upstream, and the further from upstream you get the harder it is to do that work.
Btrfs has to be rebased _every_ release. If moves too fast and there is so much work being done that you can't just cherry pick individual fixes. This makes it a huge pain in the ass.
Then you have RHEL's "if we ship it we support it" mantra. Every release you have something that is more Frankenstein-y than it was before, and you run more of a risk of shit going horribly wrong. That's a huge liability for an engineering team that has 0 upstream btrfs contributors.
The entire local file system group are xfs developers. Nobody has done serious btrfs work at Red Hat since I left (with a slight exception with Zach Brown for a little while.)
Suse uses it as their default and has a lot of inhouse expertise. We use it in a variety of ways inside Facebook. It's getting faster and more stable, admittedly slower than I'd like, but we are getting there. This announcement from Red Hat is purely a reflection of Red Hat's engineering expertise and the way they ship kernels, and not an indictment of Btrfs itself.
I imagine they pay significantly less than the other companies (e.g. Facebook) who want to hire Btrfs devs can afford to, too.
As a result, it is not at all surprising that Red Hat ends up functioning as somewhat like a baseball farm team for companies like Facebook, Google, etc. who are willing to pay more and have more liberal travel policies than Red Hat. If someone can become a strong open source contributor while working at Red Hat, they can probably get a pay raise going somewhere else.
There is a trade off --- companies that pay you much more also tend to expect that you will add a corresponding amount of value to the company's bottom line. So you might have slightly more control over what you choose to work on at Red Hat.
Fragmentation is also an issue, and xfs_fsr should be run at regular intervals to "defrag" an XFS file system. (I assume that) BtrFS handles this more intelligently.
I'd love to see XFS get some or all of these features.
Stop backporting fixes. You're forking the codebase.
Ship exactly what upstream provides.
Teach upstream projects how to do better release engineering if they're abandoning major releases to early or breaking API/ABI in a minor release.
Stop backporting fixes. You're forking the codebase.
edit: also stop incorrectly backporting security fixes and creating new CVEs. Seriously. Stop it.
Not all upstreams are interested in doing release engineering. There are non zero costs to doing it. It can eat up time that can be spent on bug fixes and features, or even make it too costly to change direction if a certain approach to implementation is proving more difficult than it should be.
Look at the Linux kernel. The only reason there is a stable kernel series is because Greg K-H decided it was important enough. He was unable to convince any other developers to go along with it, and eventually the decision was "if you want to support it, then you can do it."
Do you consider the stable kernel series a fork of the codebase? Should everyone be running the newest kernel every release despite the plenty of regressions that appear?
Kernel developers are not interested in making every change in such a slow and controlled manner as to avoid any regressions. And it works for them. They get a lot of stuff done, and come back and fix the regressions later.
[0]: https://fedoraproject.org/wiki/Staying_close_to_upstream_pro...
"edit: also stop incorrectly backporting security fixes and creating new CVEs. Seriously. Stop it."
Can you give some examples of cases where Red Hat introduced bugs in their backported patches? I follow RHEL CVEs relatively closely (because some of my packages are derived from their packages), and I can't think of an example of that happening. Debian has done so, but very rarely, that I can recall. (And, Ubuntu, too, since they just copy Debian for huge swaths of the OS.)
If you don't care that much about development velocity, it's really easy to make something that is super stable.
If you only care about making things work on a very narrow use cases (to support the back end of a particular company's web servers, or just to support a single embedded device), life also gets much easier.
If you want to "move fast and break things", that's also viable.
Finally, if you have unlimited amounts of head count, life also becomes simpler.
Different parts of the Linux ecosystem have different weights on all of these issues. Some environments care about stability, but they really don't care about advanced features, at least if stability/security might be threatened. Others are interested in adding new features into the kernel because that's how they add differentiators against their competitors. Still others care about making a kernel that only works on a particular ARM SOC, and to hell if the kernel even builds for any other architecture. And Red Hat does not have infinite amounts of cash, so they have to prioritize what they support.
So a statement such as "Teach upstream projects how to do better release engineering", is positively Trumpian in its naivete. Who do you think is going to staff all of this release engineering effort? Who is going to pay for it? Upstream projects consists of some number of hobbists, and some number of engineers from companies that have their own agendas. Some of those engineers might only care about making things better for Qualcomm SOC's, and to hell with everyone else. Others might primarily interested in how Linux works on IBM Mainframes. If there are no tradeoffs, then people might mind work that doesn't hurt their interests, but helps someone else. They might even contribute a bit to helping others, in the hopes that they will help their use case. That's the whole basis of the open source methology.
But at the same time you can't assume that someone will spend vast amounts of release engineering effort if it doesn't benefit them or their company. Things just don't work that way. And an API/ABI that must be stable might get in the way of adding some new feature which is critically important to some startup which is funding another kernel engineer.
There is a reason why the Linux ecosystem is the way it is. Saying "stop it" is about as intelligent as saying that someone who is working two 25 hour part-time jobs should be given a "choice" about her healthcare plan, when none of the "choices" are affordable.
Upstreams don't have the resources to do proper release engineering, they're busy working on new features. The fact that SUSE and Red Hat spawned from a requirement for release engineering that upstreams were not able to provide should show that it takes a lot more work than you might think.
Also, can we please all agree as a community that writing patches and forking of codebases is literally the whole point of free software? If nobody should ever fork a codebase then why do we even have freedom #1 and #2? The trend of free software projects to have an anti-backport stance is getting ridiculous. If you don't want us to backport stuff, stop forcing us to do your release engineering for you.
Honestly what I think we need is to have containers that actually overlay on the host system and only include whatever specialised stuff they need on top of the host. So updates to the host do propagate into containers -- and for bonus points the container metadata can still be understood by the host.
Again and again we see that without any financial incentive, developers are loath to put any effort into backwards compatibility and interface stability.
At the same time they all want people to be running their latest and shiniest.
So in the end, what will happen is that each "app" will bundle the world, or at least as much as they feel they need to.
For many customers of Red Hat, that mantra is the very reason they use RHEL in the first place.
Sadly i feel that more and more upstream wants it both ways, be able to push their latest and shiniest, and keep ignoring any need for interface stability etc.
Frankly i suspect the end result of the likes of Flatpak will be that upstream push whole distros worth of bundled libs, just so they don't have to consider interface stability as they pound out their shinies in their best "move fast and break things" manner.
The number of times the APIs changes from under your feet is astounding - even with just keeping up with Red Hat, we spend around 4-6 times the engineering time on the driver compared to what we do with the Windows version of the driver, tracking upstream gave us almost an order of magnitude more work (And keep in mind that /only/ supporting the most recent upstream kernel is rarely an option - several versions need to be supported concurrently)
From the perspective of legacy systems, Red Hat's approach is more comfortable.
It's a clear indicator that RedHat doesn't want or can't support btrfs.
Which is a reflection of btrfs AND RedHat: the effort required to maintain it, the lack of usage in RHEL paying customers, the immaturity/fast development of the filesystem.
I'm pretty sure, had RH wanted they could either hire or assign engineers to maintain the btrfs code, take care of patches from upstream, etc. So why didn't that happen? I wonder what is your opinion on that.
I see a bunch of possibilities (not necessarily independent ones):
1) Politics. Perhaps RH wants to kill btrfs for some reason?
I see this as rather unlikely, as RH does not have a competing solution (unlike in the Jigsaw controversy, where they have incentives to kill it in favor of the JBoss module system).
2) Inability to hire enough engineers familiar with btrfs, or assign existing engineers.
Perhaps the number of engineers would be too high, increasing costs. Especially if not only to maintain the RHEL kernels, but to contribute to btrfs and move it forward.
Or maybe there's a pushback from the current filesystems team, where most people are xfs developers?
3) Incompatible development models.
If each release requires a rebase, perhaps supporting btrfs would require too much work / too many engineers, increasing costs? I wonder what Suse and others are doing differently, except for having in-house btrfs developers.
4) Lack of trust btrfs will get mature enough for RHEL soon.
It may work for certain deployments, but for RHEL customers that may not be sufficient. That probably requires a filesystem performing well for a wider range of workloads.
5) Lack of interest from paying RHEL customers.
Many of our customers have RHEL systems (or CentOS / Scientific Linux), and I don't remember a single one of them using btrfs or planning to do so. We only deal with database servers, which is a very narrow segment of the market, and fairly conservative one when it comes to filesystems.
But overall, if customers are not interested in a feature, it's merely a pragmatic business decision not to spend money on it.
6) Better alternatives available.
I'm not aware of one, although "ZFS on Linux" is getting much better.
So I tend to see this as a pragmatic business decision, based on customer interest in btrfs on RHEL vs. costs of supporting it.
This move by Red Hat must be seen as a provocation of Oracle, to force either greater cooperation and compliance in producing a stable BtrFS for RHEL, or the release of ZFS under a compatible license. Red Hat has put an end to BtrFS for now, and Oracle will have to go to greater lengths to use it in their clone. Customers also will not want it if it does not run equally well between RHEL and Oracle Linux.
It is obvious that Oracle will have to assume higher costs and support if they want BtrFS in RHEL. Red Hat is certainly justified in bringing Oracle to heel.
Oracle recently committed preliminary dedup support for XFS, so they must be intimately aware of the technical and legal issues behind Red Hat's move.
https://blogs.oracle.com/linuxkernel/upcoming-xfs-work-in-li...
> This move by Red Hat must be seen as a provocation of Oracle
I doubt it.
https://en.wikipedia.org/wiki/Btrfs
"... initially designed at Oracle Corporation for use in Linux."
More likely they'll support it only on their Linux.
https://docs.oracle.com/cd/E37670_01/E57668/html/ol_kern_65r...
Assuming that RHEL v8 strips BtrFS, Oracle's RHCK will have to add support back in, and thus no longer be "compatible." Without that support, some filesystems will fail to mount at boot. In-place upgrades from v7 to v8 will be problematic.
Oracle has worked very hard to maintain "compatibility" with Red Hat, even going so far as to accept MariaDB over MySQL. Their reaction to the latest "poison pill" will be interesting.
An in-place upgrade from v7 to v8 could easily get hosed.
Now as to
> "Why Red Hat does not have engineers to support btrfs?"
You have to understand how most kernel teams work across all companies. Kernel engineers work on what they want to work on, and companies hire the people working on the thing the company cares about to make sure they get their changes in.
This means that the engineers have 95% of the power. Sure you can tell your kernel developer to go work on something else, but if they don't want to do that they'll just go to a different company that will let them work on what they care about.
This gives Red Hat 2 options. One is they hire existing Btrfs developers to come help do the work. That's unlikely to happen unless they get one of the new contributors, as all of the seasoned developers are not likely to move. The second is to develop the talent in-house. But again we're back at that "it's hard to tell kernel engineers what to do" problem. If nobody wants to work on it then there's not going to be anybody that will do it.
And then there's the fact that Red Hat really does rely on the community to do the bulk of the heavy lifting for a lot of areas. BPF is a great example of this, cgroups is another good example.
Btrfs isn't ready for Red Hat's customer base, nobody who works on Btrfs will deny that fact. Does it make sense for Red Hat to pay a bunch of people to make things go faster when the community is doing the work at no cost to Red Hat?
Release under a compatible license would likely see a ZFS kernel module appear in EPEL immediately; Red Hat would likely replace XFS with ZFS as the default in RHEL8 were this legally possible.
Oracle supports BtrFS in their Linux clone of RHEL. It certainly appears that Red Hat is swallowing a "poison pill" to increase Oracle's support costs (and I'm surprised that they have not swallowed more).
http://docs.oracle.com/cd/E52668_01/E54669/html/ol7-about-bt...
With these new added costs, Oracle might find it cheaper to simply support the code for the whole ecosystem (CentOS and Scientific Linux included). Given the adversarial relationship that has developed between the two protagonists, an enforceable legal agreement would likely be Red Hat's precondition.
Otherwise, BtrFS has been mortally wounded.
FWIW I haven't said anything about Oracle & btrfs ...
>> "Why Red Hat does not have engineers to support btrfs?"
> You have to understand how most kernel teams work across all companies. Kernel engineers work on what they want to work on, and companies hire the people working on the thing the company cares about to make sure they get their changes in.
> This means that the engineers have 95% of the power. Sure you can tell your kernel developer to go work on something else, but if they don't want to do that they'll just go to a different > company that will let them work on what they care about.
> This gives Red Hat 2 options. One is they hire existing Btrfs developers to come help do the work. That's unlikely to happen unless they get one of the new contributors, as all of the seasoned developers are not likely to move. The second is to develop the talent in-house. But again we're back at that "it's hard to tell kernel engineers what to do" problem. If nobody wants to work on it then there's not going to be anybody that will do it.
Sure, I understand many developers have their favorite area of development, and move to companies that will allow them to work on it. But surely some developers are willing to switch fields and start working on new challenges, and then there are new developers, of course. So it's not like the number of btrfs developers can't grow. It may take time to build the team, but they had several years to do that. Yet it didn't happen.
> And then there's the fact that Red Hat really does rely on the community to do the bulk of the heavy lifting for a lot of areas. BPF is a great example of this, cgroups is another good example.
I tend to see deprecation as the last state before removal of a feature. If that's the case, I don't see how community doing the heavy lifting makes any difference for btrfs in RH.
Or are you suggesting they may add it back once it gets ready for them? That's possible, but the truth is if btrfs is missing in RHEL (and derived distributions), that's a lot of users.
I don't know what are the development statistics, but if majority of btrfs developers works for Facebook (for example), I suppose they are working on improving areas important for Facebook. Some of that will overlap with use cases of RHEL users, some of it will be specific. So it likely means a slower pace of improvements relevant to RHEL users.
> Btrfs isn't ready for Red Hat's customer base, nobody who works on Btrfs will deny that fact. Does it make sense for Red Hat to pay a bunch of people to make things go faster when the community is doing the work at no cost to Red Hat?
The question is, how could it get ready for Red Hat's customer base, when there are no RH engineers working on it? Also, I assume the in-house developers are not there only to work on btrfs improvements, but also to investigate issues reported by customers. That's something you can't offload to the community.
I still think RH simply made a business decision, along the lines:
1) The btrfs possibly matters to X% of our paying customers, and some of them might leave if we deprecate it, costing us $Y.
2) In-house team of btrfs developers who would work on it and provide support to customers would cost $Z.
If $Y < $Z, deprecate btrfs.
Oracle owns a lot of patents and I suspect both ZFS and BtrFS rely on some.
Perhaps Redhat could help to develop snapshots on XFS. It's not the only feature XFS is missing, but it's a start.
Upstream has been working on adding more btrfs-like features to XFS, but I believe that RHEL encourages using devicemapper snapshots (which you then format with XFS).
Exactly, mountable and mergable snapshots have been supported by a LVM/devicemapper stack for a long time.
The only connection Oracle has to ZFS on Linux is ownership of some patents that the license allows you to use, their reluctance is based on distribution issues between the GPL and CDDL.
Any Linux contributor could also try to enforce it, which is why the license incompatibility is the issue stopping them. Oracle holds no special power.
Most Linux contributors want Linux to succeed. I don't think it's at all clear that corporate Oracle prefers Linux to succeed - at least not if higher adoption of Solaris is an alternative.
[1] (I guess IBM and Microsoft come to mind... but they don't have any special investment in ZFS)
The license issue is what's keeping Red Hat from using ZFS, not some rivalry with Oracle.
The only issue I'm aware is that mixing the two would violate GPL.
GPL says: Source code must be licensed under GPL.
If you follow the conditions of GPL, you are violating the condition of the CDDL. If you are following the conditions of CDDL, you are violating the GPL. Basic binary logic.
To add: "the engineers who had written the Solaris kernel requested that the license of OpenSolaris be GPL-incompatible". A license is really just an written intention of the author on what conditions copyright law restrictions may be legally ignored. In this case, those wishes had a very explicit intention. However those using the license today has had a general change of heart, and those with GPL interest has a general stance that no FOSS project will ever sue an other FOSS project over license incompatibility. As such, the risk of lawsuit is really just a company suing an other company under the technicality of incompatibility.
Naturally some organizations won't intentionally break copyright law just because no one will sue.
>If you are following the conditions of CDDL, you are violating the GPL. Basic binary logic.
Relationship between licenses can be transitive but not commutative.
As far as I know CDDL allows using with code under GPL but GPL does not allow using code under CDDL. CDDL copyright owners have no case, GPL copyright owners have.
The question is: If I'm incorrect, what in CDDL prevents using with GPL?
CDDL has this text: "Any Covered Software that You distribute or otherwise make available in Executable form must also be made available in Source Code form and that Source Code form must be distributed only under the terms of this License"
So you take some CDDL code, and some GPL code, and you put that whole new source code tree under GPL in order to fullfill the GPL license condition. Are you then in compliance with the CDDL code? My concussion is that you are not, as that would be in conflict with the above condition of the CDDL. The source code tree would not be "distributed only under the terms of this license".
I take a CDDL licensed source code file. I take a GPL licensed source code file. I add inline the GPL licensed code to the CDDL licensed file, and release an executable form of the result. In order to comply with the GPL I then give out a single source code file under the GPL license terms with the code from the two files.
Is this in compliance with the CDDL terms and conditions?
https://opensource.stackexchange.com/questions/2094/are-cddl...
https://github.com/zfsonlinux/zfs/blob/master/OPENSOLARIS.LI...
Also, the SO answer mentions consumer protection laws - but AFAIK they generally only apply to consumers - not businesses. So the GPL 0 clause might be void in many jurisdictions for individuals but still valid for businesses.
The GPL on the other hand is a strong copy left. If you link against GPL code, your code must also be licensed as GPL.
This means the Linux copyright owners could sue the distributers of ZoL binaries, but Oracle could not.
Oracle has the power to allow their ZFS code to be relicensed as GPL, removing this road block, but they have no incentive to do so.
Given they are discontinuing Solaris and all-in on Red Hat Enterprise Linux I can't help but wonder why they don't do more with ZFS on Linux and therefor wonder if the NetApp patent suits or some other patent suit is preventing them from doing anything in the background.
Many people don't realise that these crappy patent suits in the background prevent all sorts of really basic stuff, like the fact most things now bounce through a cloud server (like Facetime) because there's a patent troll for peer-to-peer communications. And it's causing total waste as a result :( It also seems likely that prevented facetime becoming an open standard as Apple original promised. This is only 1 example though.
Oracle is NOT discontinuing Solaris. This FUD must die.
Killing OpenSolaris, and talking up SPARC so much, made people think that a) Larry just wants to vendor lock them, b) doesn't care about x86 support because it makes vendor lock-in harder for Oracle, c) the OpenSolaris derivative community will not be able to compete with Linux. So everyone has grudgingly accepted that Linux is it for the enterprise Unix market.
I hate this as much as you do. I <3 Solaris/Illumos. Illumos derivatives have their niches, no doubt, and I want to be able to use them much more. But that's not how business people think.
I'm not sure that Oracle could turn this impression around at this point. To begin with it would have to restart OpenSolaris, and that might not be enough. OpenSolaris greatly helped Sun overcome resistance to Solaris, but it only went so far, so Oracle will have to do even more work to make Solaris' future bright.
This blog post is as relevant today as ever: https://blogs.oracle.com/bmc/the-economics-of-software
(And yes, it's STILL hosted at blogs.oracle.com. I'm almost afraid of mentioning it: who knows, it might get removed if Oracle execs notice it.)
https://theregister.co.uk/2017/01/18/solaris_12_disappears_f...
XFS has outperformed EXT4 in almost all "high" use-cases in my experience and testing: Large files (500GB~) or many small files (128k files * 2,400,000 or so). EXT4 under those loads is comically bad.
BTRFS is also terrible at this, only XFS and ZFS are good at handling it.
Here is one benchmark, but I have seen plenty of similar benchmark results for PostgreSQL showing the same thing: https://blog.pgaddict.com/posts/postgresql-performance-on-ex...
There are various guides around for tuning ZFS and database servers to try reduce that duplication, for example you can disable the InnoDB double write buffer because ZFS guarantees you don't need it. You also need to tune recordsize to match the database page size so that you don't accidentally create large multi page blocks.
So if you only need a plain filesystem, ext4/xfs are great and you will get better performance.
If you need/want snapshots, e.g. to do backups that way, it makes sense to look at ZFS.
Also: I love how that comment sits at -4, as if downvoting it will somehow discredit the data point.
> Creating a new directory entry on an idle machine with plenty of CPU and memory takes seconds, ditto deletions.
I think your answer lies in your premise then. It's not representative.
I've been using XFS for 10 years without the issues you seem to be having.
I guess the general stuff is: the easy default partitioning setup you get from a Linux distro is total bs, you need more RAM than you think you do, the way you're serving files or accessing the system (NFS!) has plenty of ways to screw things up as well, and tens or hundreds of millions of files is not any filesystem's ideal use case. The classic IRIX workload would be guaranteed-rate streaming of large media files, and the Linux port of the filesystem obviously inherited a lot of that system's traits (without the GRIO).
XFS has received some very serious performance improvements in the past couple of years to address indexing, large volumes of metadata, and so on, so that'd be one very relevant thing. Dave Chinner's talks are worth the time to watch if you're interested. You would be giving bad advice if you steered people one way or the other with regard to filesystems based on a seven-year old project (unless you've refreshed that system much more recently, of course).
That's probably the difference right there. Thanks for pointing that out.
[1] https://access.redhat.com/documentation/en-US/Red_Hat_Enterp...
xfs_repair will complain if there is a journal present and tell you to mount the fs to replay the journal. But mount would refuse, saying the fs was inconsistent. So the only option was to xfs_repair -L to just throw out the journal.
Then, xfs_repair sucked up something like 30GB or more of RAM, so I had to make a huge swapfile so that the kernel would OOM the repair.
Then, after roughly 20 to 30 hours of repair it would exit with an error. At that point it would actually mount, but hitting certain areas of the filesystem would trigger the inconsistency again and start the entire process over.
In the end I couldn't fix it and sadly had to reformat. I chose ext4 when I did—I've had lots of experience with ext3 and 4 and I've never had a filesystem that I couldn't at least make consistent again (even if it loses some data).
I consider checksumming important. Do others? What is the solution? What other file systems offer that sort of capability?
Snapshotting is a second go-to function. Particularly when it is integrated into the LXC container creation process. (There was a comment elsewhere here which said LXC is on it's way out.... huh? what?)
Do you have a source for this? So far I believed that bit-rot rates are pretty similar.
Additionally use smartmontools and configure it to do a short self test each night, and a long self test (i.e. full disk read) each week.
This will catch/flag errors early, which mdadm will then detect.
On disk B this file has rotted.
When reading the file SomeFile into memory, the read will be distributed among the disks (for performance reasons) (and it will probby need to span a multiple of the stripe size).
Ok, file is read into memory, including the bitrotted part from disk B. Now we write the file blocks back - as one does.
Voila! Both disks now contain the bitrot. And mdadm will not complain - disk A and B are identical for the area of file SomeFile.
Just use ZFS. Even on a single disk setup you will at least not get silent bit rot.
- adage cited in the Mythical Man Month
Are you referring to mirroring a volume or dataset on a single disk? Why would you want to do that instead of mirroring among multiple drives?
two sets of say 5 disks in a mirror raidz1 would still fail if a disk in one set failed and a disk in the other set failed. I guess you could do a stripe setup of 5 sets of 2 disks in mirrors. Still it seems wicked risky to me. I do agree though mirroring has been the best for speed but a lot of that changes with nicer SSDs especially NVMe ones.
I'd setup a large pool with mirror vdevs, i.e. n sets of 2 disks per mirror.
My half-remembered reasoning was that backups manage the risk you'll lose data. But replacing a disk in a mirror vdev is much easier, and faster, than doing so with RAIDZ.
The risk of RAIDZ is that resilvering impacts multiple vdevs, is much more intensive than a simple mirror resilvering, and thus the probability that additional drives will fail is much higher.
Here's a blog post that I definitely read the last time I was reading up on this:
- [ZFS: You should use mirror vdevs, not RAIDZ. – JRS Systems: the blog](http://jrs-s.net/2015/02/06/zfs-you-should-use-mirror-vdevs-...)
- [You should use mirror vdevs, not RAIDZ. : DataHoarder](https://www.reddit.com/r/DataHoarder/comments/2v0quc/you_sho...)
Moreover, if it doesn't always read both copies of the data (which it may well not, for performance reasons), then you have the possibility of silently propagating damaged data to all mirrors in the case that damaged data is returned to an application and the application then rewrites said data.
Compare that to a filesystem with checksums, which, in addition to being able to detect such a problem, could also continue to function completely correctly in the face of it.
Mirrors and RAID5: there's obviously no way that `md` software RAID can help, since it doesn't know which is correct. What about RAID6 though? Double parity means `md` would have enough information to determine which disk has provided incorrect data. Surely it does this, right?
Wrong. In the event of any parity mismatch, `md` assumes the data disks are correct and rewrites the parity to match. See "Scrubbing and Mismatches" section in `man 4 md`:
https://linux.die.net/man/4/md
If you scrub a RAID 6 array with a disk that returns bad data, `md` helpfully overwrites your two disks of redundancy in order to agree with the one disk that's wrong. Array consistent, job done, data... eaten.
Any recommendations for detecting/correcting bitrot with RHEL 7.4 at the filesystem or lower levels?
Depending on the requirements of your use case different choices will make more sense. It's important to remember that RHEL is used for enterprise customers, and what might be common in the enterprise world might not be common for yours, and vice versa. Certainly, if you are using a cluster file system, it makes no sense to do checksum protections at the disk file system level, because you will be using some kind of erasure coding (e.g., Reed Solomon error correcting codes) to protect against node failure. This will also take care of bit flips.
If you are using cloud VM's, or if you are using Docker / Kubernetes, then LXC won't make sense. It all depends on your technology choices, and so it's important to look at the big picture, not just at the individual file system's features.
> this target do not provide error correction, only detection of error (such a tool could be written on top of dm-integrity though)
It does however return an error if the integrity check fails, so if you put mdadm on top, mdadm can repair the erroneous block. I've tested this and am currently running it on a 32TB array.
So I moved back safely to ext4 and never looked back!
Nowadays, i must say that i very much prefer a stable filesystem with as little complicated logic as possible. I actually never use snapshots or subtrees! I never put another disk in my laptop (where would that go?!) so i don't need to do dynamic resizing (while online of course!). All this makes the filesystems a lot more complex then it has to be. I've also run into problems using ZFS on Solaris some years back which took ~2 weeks in dtracing what the hell is going on. Of course it was related to CoW.
My lessons learned: Check your requirements. Will you really need and use subtrees/snapshots/XYZ on your system? Will you really need to do online-resizing? If not, just use a stable, simple filesystem. There are perfect usecases for ZFS or btrfs. But not everyone needs the advanced features.
It's a valid question, but not the best one. Almost nobody needs snapshots. But they make things easier. You most likely don't need a journaled fs in your laptop either (battery level notification should take care of the issues). But it does make life better.
"Need" is not the threshold I'm interested in. Most features, I'd like. One feature I think I do need most is scrubbing, which is still absent from most filesystems :(
But that's only me. Your experience may differ very much :)
Well, many of us have experienced a botched system package upgrade or two. If the file system supports snapshots, then the package manager could automatically ensure fully atomic package upgrades.
That should be reason enough, I should think.
Re: The data loss issue: Yes, I've actually have XFS completely throw away a file system upon a hard power-off + boot-up cycle. (This was ages ago, I'm sure it's improved heaps since then.)
That's exactly what openSUSE / SLE do with snapper. Every upgrade or package install with YaST/zypper creates two snapshots (before/after) and you can easily rollback to an older snapshot (even doing so from GRUB). This has been enabled by default for years.
The atomic updating is a very interesting topic and the reason why i find ostree/guix/nixos very appealing. Note that neither ostree nor guix or nixos make use of filesystem snapshots, afaik. OSTree even documents why it won't use filesystem snapshots: https://ostree.readthedocs.io/en/latest/manual/related-proje... Debians dpkg does not use snapshots as well.
So, it's a definitely a nice-to-have, but not something i need, because i can handle dpkg/apt much better then i could handle filesystem internals.
That's sort of the point: I don't want the complexity in the filesystem, but i am fine with it in userspace. I can use snapshots on filesystem level. Or i can use other backup tools in userspace. While it's certainly neat that the filesystem can do that, i'm perfectly fine with handling backups on another level.
Another example: It's certainly neat that there are a bunch of distributed filesystems (which by the way have A LOT of complexity and often can't handle all workloads you would expect from a filesystem). But i'd rather use either an S3-like network storage or build a system that scales well without relying on Ceph/Gluster/Quobyte/etc.
For example, in a hypothetical distributed system i'd rather use Cassandra and distribute data over commodity hardware then use Ceph. I'd rather handle problems with data persistence/replication on the cassandra level then debugging on file system level. Especially, when Cassandra has a problem i'll most likely be able to access all data atleast on the filesystem level. When my filesystem is borked, i'm in a much worse situation.
I don't think you're seeing my point. You wouldn't have do anything -- it would all be done automatically as long as your file system supports snapshots.
BTW, to your "I know how to use dpkg/apt": It's not about knowledge. I could well be said to be at an "advanced" level of expertise in system maintenance, but "system upgrade" fuckups had nothing to do with me, but everything to with bad packaging and/or weird circumstances such as a dist-upgrade failing midway through because some idiot cut a cable somewhere in my neighborhood.
While Nix and the like are nice and all, they're currently suffering from a distinct lack of manpower relative to the major distributions. They also don't quite fully solve the "atomic update" problem, but that's a tangent. Then, OTOH, some of them have other advantages such as the easy of maintaining your full system config in e.g. Git. Swings and roundabouts on that front. FS support for snapshots would help everybody.
Did you read the link from the ostree people? Let's pretend Debian 10 offers to choose between OStree-like updates and btrfs snapshots: I'd probably choose OStree and stick to ext4/xfs.
Yes, but it SHOULD, just because ALL REASONABLE FILE SYSTEMS SHOULD SUPPORT SNAPSHOTS. Therefore dpkg should assume that such support is avaiable, or at the very least take advantage of it, when available.
Just to reiterate: You (impersonal!), the "ignorant user", shouldn't have to even have to think about it.
Does this make my point clear?
(I'm only being this obtuse because you're saying "you're absolutely right", but apparently not seeing my point. I'm assuming it's some form of miscommunication, but it's difficult to tell.)
EDIT: Hehe, I'm sorry, that sounded much more aggressive than I intended. I just think that us software developers could and should(!) do much better by our users than we(!) currently do. My excuse is that most of my stuff is web-only, so at least I can't do the accidental equivalent of "rm -rf /", but...
Another use for snapshots is backups. I love `zfs send` - it makes backup braindead-simple.
Here is one way to do simple, secure scrubbing on Linux without any intrusive system changes. It is mildly restrictive, but works.
First, you need a small, dedicated partition, but it only needs to be around 16MB or so. Resizing an existing partition down (tune2fs will happily resize a mounted ext4 filesystem, but you'll probably still need to reboot to reload the partition table once you've resized that too) will give you a bit of space.
Now you have a small area of the disk that occupies a known range of sectors, and because you have no TRIM, writes to this area will be properly deterministic. Good.
Create and mount a new filesystem without a journal on the new partition. ext2 could work here (:D), you could `mkfs.ext{3,4} -O ^has_journal`, or you could use filesystem defaults and simply overwrite the entire partition with /dev/urandom later.
Make a sparse file with fallocate (make sure the file system you create the file on can handle sparse files) that is big enough to handle the biggest file.
Create a LUKS volume with a detached header inside the new sparse file, and store the detached header metadata into a file in the new journal-less partition.
Create an ordinary filesystem inside the LUKS volume.
Now you have a Rube Goldberg sparse file. You've moved the deterministic-writing/journal-less stage into a tiny key, which is a lot easier to manage than a whole gigabytes+-large partition.
As an alternative you could drop the key onto a flash drive, and nuke the flash drive when you wanted to kill the data. That's kind of wasteful though (and it carries the same flash-drive-quality risks as copying the only copy of the data itself onto the flash drive).
LUKS was designed such that if you lose the key(s) or the detached header, all that's left is statistically random garbage.
Scrubbing means to read all the data off a filesystem and compare it against its checksums, so that you are confident nothing has happened to the data (hardware failures, cosmic rays, whatever).
ZFS and btrfs have specific scrub commands that do that.
There's no scrubbing available for a system which does not keep some form of checksum/crc/hash of the data.
I think that you are talking about secure delete procedures.
I actually tried to delete this comment for unrelated reasons shortly after posting it, but was unable to. Now I feel doubly stupid.
I also really like that I don't need to partition my disk, if it turns out that /tmp needs > 10% of the disk for whatever reason: no problem!
And as I like my data, I appreciate checksumming and copy-on-write.
I haven't noticed any bad slowdowns compared to ext4 on my Debian laptop I used before.
Did you know you can 'zfs send' snapshots to rsync.net ?[1][2]
[1] https://arstechnica.com/information-technology/2015/12/rsync...
In general I feel that the more you go into "enterprise" storage levels, the more XFS pulls ahead from the EXT family. i.e. laptops and small servers are not where the difference lies.
I don't get why apt syncs so often - isn't the main point of log-structured file systems their ability to recover after a crash or powerloss? If so, why should you need to sync more than one every ten seconds or so?
When you move to a more advanced setup such as ZFS clones, you could do the full upgrade with a cloned snapshot, and swap it with the original once the changes were complete. This would avoid the need for all intermediate syncs--if there's a problem, you can simply restart from the starting point and throw all the intermediate state away.
I've had older rpm & yum/dnf failures multiple times leaving me in weird inconsistent states from crashes or power losses etc. not conclusive but anecdotal experience - It's also possible it's been improved.
Meanwhile you can disable the file syncing with the apt preference dpkg::unsafe-io (google will be required for the exact syntax and file in /etc/apt - fairly sure you can cmdline it also)
oracle pays the developers of btrfs [0]
redhat hates the guts of oracle, since oracle released oracle linux, which is a clone of redhat enterprise (based on centos)
so, redhat wants to cripple btrfs and hurt oracle.
However, btrfs is my favorite FS, been using it on my home computer and backup drives for at least 6 years, before it was included in the kernel, love the subvolumes, snapshots, and compression; never had issues with it .
[0] https://oss.oracle.com/~mason/
[Update] Chris mason no longer at Oracle since 2012
https://en.wikipedia.org/wiki/Btrfs#History
"In June 2012, Chris Mason left Oracle for Fusion-io, which he left a year later with Josef Bacik to join Facebook; while at both companies, Mason continued his work on Btrfs."
* Liu Bo * Anand Jain
Slightly off topic. I chose btrfs as my main filesystem recently on a system running Ubuntu/Xubuntu. I have done some research on backing up (with the advantage of snapshots) but it looks like there aren't (m)any graphical tools (this gets a little more confusing with /@ and /@home subvolumes on the same partition being treated separately for snapshots, AFAIK).
Do you manage it all from the command line and/or do you have any suggestions for graphical tools to do "as-is clones of entire partitions" (and also incremental backups) to local external drives (not over the network)? Or if you could point to any great documentation or blog posts on this topic, that'd be helpful too (I have read some bits of the btrfs wiki and the btrfs parts in the Arch wiki).
Currently I'm doing a plain rsync using Grsync, and not really taking advantage of btrfs features like snapshots.
The main reason I'm looking at avoiding the command line is to make it easier for others around me to use it.
btrfs subvolume snapshot /source/drive/folder/ /source/drive/folder/.snapshots/snapshot-`date +"%Y-%m-%d-at-%I-%M%P"`
this will create a snapshot with date and time attached to the snapshot name
you can more info here
https://btrfs.wiki.kernel.org/index.php/SysadminGuide#Managi...
make sure to remember to delete old snapshots, otherwise you'll run out of disk space and not know where it went
I had auto hourly snapshots and sometimes when it deleted one my entire system would hang for a few seconds and occasionally 10s of seconds.
Having said that I do suspect that might be partially related to also using ecryptfs on top, but still.
I think there are solid technical reasons to discourage Btrfs use, just to quote from the official wiki [0]:
> The parity RAID code has multiple serious data-loss bugs in it. It should not be used for anything other than testing purposes.
Now I don't know if this issue has been addressed already, or which kernels are affected, but the fact that there is a prominent warning on the wiki speaks for itself.
Personally, I'm a happy btrfs user deploying a mixed-disk-size array without parity, with the hope to add redundancy some time in the future. Currently, btrfs is the only FS allowing to mix disks of any size and to run an optimal configuration on top of them [1].
The reason ZFS doesn't support it, and absolutely 0 enterprise storage devices support this is because as the disks fill up, you sacrifice both performance and redundancy. Synology won't even support it on their high-end devices for this very reason. They'll only do it on their devices targeted at home use.
mdadm will give you 1 TB in RAID 1, or 1.5 TB in RAID 10 (constrained by the smallest drive).
btrfs will give you 3 TB in RAID 1 (constrained by the sum of the smallest drives).
btrfs also allows per-subvolume raid policies. So you could, for example, give users an "archive" subvolume in their home directory. You could then mark this as RAID 1 or RAID 5 (because you don't care so much about performance) while the main /home filesystem is RAID 10.
Unfortunately the RAID code is all horribly broken.
My bad.
[1]: https://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg...
You literally have no idea what you're talking about, and I doubt you've used btrfs seriously, or you wouldn't talk this shit. The fact it's been upvoted so heavily just shows what absolute technically-false nonsense will draw support at HN.
I've finally broken and installed ZOL (ZFS on Linux) after trying btrfs repeatedly over the last three years. ZFS is already a breath of fresh air and I've only been using it a couple of months. For whatever reason, btrfs came together as a messy hodge-podge, and it shows in bad performance for many use cases (e.g. "omg I forgot nodatacow"), buggy implementations, difficult user interfaces, kernel bugs, etc.
btrfs needs a reboot (I hear bcache? is trying). Meanwhile, everyone should stop getting hung up on the arcane licensing details and just use ZFS directly. It can't be distributed as part of the kernel, but that's why we have distributions, isn't it? They bundle all that crap together for us. There shouldn't even be the normal OSS infighting because this isn't a proprietary blob or something, it's just using a license that's GPL-incompatible.
The best thing Linus could do for the community at large would be to fork and start committing to ZOL, giving it a tacit endorsement.
It seems a bit here-say-ish though, so please don't assume I'm entirely correct on all fronts there and I'd encourage you to research it further!
On that note, their Patreon is a better primer than the website is: https://www.patreon.com/bcachefs
In other news.. consider supporting the people that support you! Personally, I spent over $100/month on Patreon. Most of those are creators rather than open source people but there are a couple of open source ones such as Ondřej Surý who works on PHP packaging in Debian/Ubuntu
I donated a lot to git-annex because he did it right, showing everywhere that you could sponsor the development.
Unfortunately, he's only getting $500/m, which isn't much even for someone with a very "off-grid" life: https://joeyh.name/blog/entry/notes_for_a_caretaker/
Discoverability is a real problem for Patreon, outside of the "most successful"
However, HN has a previous thread on it here: https://news.ycombinator.com/item?id=12410798
Second, it lacks a formally published design for these features. Because the author is the only person with knowledge of how these things might work, it makes it really difficult to mitigate the bus factor while the project is still in heavy development.
If you're going to support a code base for ~10 years, you're going to need upstream people to support it. And realistically Red Hat's comfortable putting their eggs all in the device-mapper, LVM, and XFS basket.
But, there's more: https://github.com/stratis-storage/stratisd
Btrfs has no licensing issues, but after many years of work it still has significant technical issues that may never be resolved. page 4
Stratis version 3.0 Rough ZFS feature parity. New DM features needed. Page 22 https://stratis-storage.github.io/StratisSoftwareDesign.pdf
Both of those are unqualified statements, so fair or unfair my inclination is to take the project with a grain of salt.
When I gave up on it there were also fundamental issues with metadata vs data balancing, not-really-working RAID support, and so on...
Sure, making the filesystem CoW-based means there are some inherent costs, but it allows the filesystem to implement some interesting features (e.g. snapshots) in a more efficient way. For example if you want to do snapshots with ext4/xfs, you'll probably do that using LVM (which you can see as turning the stack into a CoW). In my experience the performance impact of creating a snapshot on ext4/LVM is about 50%, so you cut the performance in half. While on ZFS the impact is mostly negligible, due to the filesystem is designed as CoW in the first place.
And thanks to ZFS we know that it's possible to implement a CoW filesystem that provides extremely stable and balanced performance. I've done a number of database-related tests (which is the workload that I do care about) and it did ~70-80% TPS compared to ext4/xfs (without snapshots). And once you create a snapshot on ext4/xfs, the performance tanks, while ZFS works just like before, thanks to the CoW design.
Unfortunately, BTRFS so far hasn't reached this level of maturity and stable performance (at least not in the workloads that I personally care about). But that has nothing to do with the filesystem being CoW, except perhaps that CoW maybe makes the design more complicated.
But as you mention, that does not say anything about CoW filesystems in general. It merely hints the BTRFS implementation in not really optimized.
FWIW while I do a lot of benchmarks (both out of curiosity and as part of my job, when evaluating customer systems), I've learned to value stability and predictability over performance. That is, if the system is 20% slower, but provides stable and predictable behavior, it's probably OK. If you really need the extra 20% you can probably get that by adding a bit more hardware, and it's cheaper than switching filesystems etc. (Sure, if you have more such systems, that changes the formula.)
With EXT4/XFS/ZFS you can get that - predictable, stable performance. With BTRFS not so much, unfortunately.
Interesting features are worthless when reading and writing data is prohibitively slow. Or when there are documented cases where updating a file in random-access manner can cause its storage requirement to balloon to blocks^2.
I would say ZFS works extremely well (at least for the workloads I care about, i.e. PostgreSQL databases, both OLTP and OLAP). I know about companies that actually migrated to FreeBSD to benefit from this, back when "ZFS on Linux" was not as good as it's today.
As to snapshots, who cares, they cost nothing to create and they do not slow down writes -- they only slow down things like zfs send (linearly) and they cost storage over time, but not much more.
Btrfs behaves basically like that with 'nodatacow' today. It will overwrite extents if there's no reflink/snapshot. If there is, CoW happens for new writes and any subsequent modifications are overwrites until there's a reflink/snapshot in which case CoW happens.
The 'nodatacow' flag can be used as either a mount option, or selectively with an xattr per subvolume, directory, or file. And in all cases, metadata writes (the file system itself) are still CoW.
Btrfs has been deprecated
The Btrfs file system has been in Technology Preview state since the initial release of Red Hat Enterprise Linux 6. Red Hat will not be moving Btrfs to a fully supported feature and it will be removed in a future major release of Red Hat Enterprise Linux.
The Btrfs file system did receive numerous updates from the upstream in Red Hat Enterprise Linux 7.4 and will remain available in the Red Hat Enterprise Linux 7 series. However, this is the last planned update to this feature.
Red Hat will continue to invest in future technologies to address the use cases of our customers, specifically those related to snapshots, compression, NVRAM, and ease of use. We encourage feedback through your Red Hat representative on features and requirements you have for file systems and storage technology.
Expensive hardware, with little gain sadly. It was a nice idea, however at its very core is a fairly large problem: converged network adaptors are problematic.
Unless you have lots of bandwidth in said adaptor (ie 56gig inifiband) you are going to get contention between network and disk IO.
What's the replacement then? ISCSI, or going back to FC? Or is everything cloud something these days? :)
Even VMWare has come on board with vSAN which pushes out vendors like EMC/NetApp because you no longer need them when you can just create it against your existing hypervisors. Sure you can run one or two less VM's on it, but you have less cost overall.
FCoE requires that the networking gear drops only one packet in 10 million or something like it, if you can make the same guarantees for iSCSI, it is for all intents and purposes the same thing. With iSCSI off-load it is even better.
iSCSI also runs across your existing network stack, and doesn't require purchasing special equipment and is better supporter across a variety of different vendors, thereby making it easier to find the gear with the features you need, rather than settling for something that supports FCoE.
Given the existing install-base of FC, I'm guessing as people upgrade to the 32/128Gbit adapters they will start to purchase disk's that can support FC-NVME as well. Which will bootstrap the market there.
Although it could go to infiniband as well, if people buy into the converged infiniband/ethernet adapter route.
To soon to tell, but a lot of it will be dependent on which technology does a better job avoiding the "forklift" upgrade problem that FCoE required.
The complexity of FCoE is staggering, and the configuration required across all the different moving pieces to make it a success made things even more difficult!
I've also been using btrfs as the backend to docker for a long time on my desktop PC and never noticed any problems. BTRFS has been rock solid for me. I don't doubt it is more unstable than other filesystems, however it seems i haven't been unlucky enough to experience any issues.
When using BTRFS, i've always stuck to the latest kernel releases, and run a scrub + balance every month. This is the advice I heard from people who used btrfs, and I wonder how many of the people who complain about data corruption do these steps. Perhaps their corruption bugs are solved in a newer kernel version. I've had multiple scrubs pick up data corruption, which other filesystems wouldn't have found.
The only time btrfs corrupted my data was when I used the ext4 to btrfs conversion tool, it created an unmountable FS and then I just migrated my data manually.
Manual balancing is a workaround for a critical flaw in the implementation.
In my last major use of Btrfs, whole archive rebuilds of Debian, it would take less than 48 hours to completely unbalance a brand new Btrfs filesystem. ~25k snapshots continuously created and deleted over the period in 20 parallel jobs absolutely toasted the filesystem, even though it was 1% utilised for the most part, 10% at peak usage.
The point I want to make is that a Btrfs filesystem can become unbalanced at some indeterminate point in the future, which makes it impossible to rely on if you want to guarantee continued service.
I've also suffered from a number of dataloss incidents which likely are fixed now, but despite lots of bugfixing, there are still major flaws to address.
[1]: https://www.freedesktop.org/software/systemd/man/systemd-nsp...
I guess one option is to pull btrfs tree into Systemd :-)
Will Redhat too (like Ubuntu) start shipping ZFS?
* Make fresh backups
* Verify the backups
* Re-install and use the backups
> Red Hat will continue to invest in future technologies to address the use cases of our customers, specifically those related to snapshots, compression, NVRAM, and ease of use.
but it's unclear what this means exactly.
I'm guessing that the ZFS licensing hairball is a bridge too far for even Red Hat, so they'll cobble together equivalent-ish functionality - even if it's not anywhere near as elegant as ZFS's integral data production and reduction.
The only potential risk is that the GPL is so virulently infectious that any driver is automatically GPL'd by virtue of its own existence as a compiled kernel module, but that possibility seems fairly remote, and it hasn't seemed to affect the distribution of other purportedly-non-GPL kernel modules.
I'm not a lawyer so maybe I'm missing something.
My guess too is that if Canonical manages to go a few years without a lawsuit from kernel copyright holders then we might see more of what it is doing. But RH would -I guess!- still suffer from patent FUD and so stay away from ZFS.
That's all fine by me. The better for RH's competition. More competition, mo' betta.
[0] https://en.wikipedia.org/wiki/Btrfs#Subvolumes_and_snapshots
With ZFS, you have a hierarchy of datasets. These inherit properties from their parents, and while the mountpoints can also mimic this hierarchy, the mountpoint property can be set independently. Btrfs couples the two concepts, forcing subvolumes to be in a specific place in the actual filesystem; zfs datasets in comparison are purely metadata and are for organisation and administration, not direct use in the filesystem hierarchy.
ZFS snapshots are read-only, and clones of these snapshots are datasets in the hierarchy. Btrfs snapshots are read-write by default, which in some ways defeats the point of a point-in-time snapshot. You can also make changes to a ZFS clone and later promote it to replace the original dataset. Likewise rollbacks. Btrfs makes no provision for doing either; you have to delete the original and then rename the snapshot, which isn't atomic. ZFS' metadata preserves all relations between datasets, snapshots and clones.
The ZFS way of doing things makes things safe and accessible for system administration. There's no way to confuse the origin of a snapshot because it's tied to a parent dataset. Likewise clones of snapshots, unless you deliberately choose to break the link. The Btrfs way looks superficially nicer, but in practice is much less flexible, and potentially more dangerous since you don't have the ability to audit what came from where and when. Btrfs snapshot performance is also abysmal. ZFS handles snapshots simply by recording the transaction ID, which makes them really lightweight (and it also provides "bookmarks" which are even lighter weight). ZFS keeps the referenced blocks in deadlists, and its performance is excellent (compare how fast snapshot deletion is between the two). ZFS also allows delegating permissions to perform snapshot, clone, rollback etc. to normal users; I'm unware of Btrfs allowing such delegation--some operations can be performed like snapshotting, but not deletion, while ZFS permits this all to be configured transparently.
>I had a unique opportunity to take a detailed look at the features missing from Linux, and felt that Btrfs was the best way to solve them.
>From other points of view, they are wildly different: file system architecture, development model, maturity, license, and host operating system, among other things
-------------------------------
>Btrfs snapshots are read-write by default, which in some ways defeats the point of a point-in-time snapshot.
Yes, and have the option of being read only for your temporal "in place" snapshots. But if I want to clone a container for instant use (as LXC or Docker does), then the RW snapshots make sense. Btrfs doesn't make a distinction between a Clone and Snapshot, they are one and the same with a flag.
> but in practice is much less flexible
Tell me more how I can mix disks of differing size in RAID on ZFS
> There's no way to confuse the origin of a snapshot because it's tied to a parent dataset
There's no confusing to the origin of my sanpshots. `btrfs subvolume list -q` shows the ancestral parent as well as the subvolume it's located in, example:
ID 6442 gen 50527 top level 751 parent_uuid 0f4442f8-6363-6944-be8d-e2b45d809352 path .snapshots/321/snapshot
> some operations can be performed like snapshotting, but not deletionSee user_subvol_rm_allowed mount option, available since Kernel 3.0
It's like comparing a car and a truck, they both have four wheels, transport passengers and cargo, and have an engine. Just because a truck runs on diesel does not make the fact that the car running on gas "wrong". Due to its fundamentally different implementation, the way the filesystem works is also different.
Yes ZFS has many more features, has been in development longer, and probably more "production ready" than BTRFS. But ZFS is not GPL compatible. And BTRFS doesn't require it's own separate cache that is apart from the normal filesystem cache.
ZFS sets a very very high bar indeed. There are things that could be done better (I've talked about some of those on HN). But pound for pound, it's the best storage stack today and has been for over a decade. ZFS is the benchmark against which all others are to be stacked. There will be applications for which you will find a more performant solution, maybe, but altogether, ZFS has been the last word in filesystems for a long time now.
The most interesting competition, IMO, is from HAMMER. We'll see how that progresses.
Not entirely. Btrfs was designed with benefit of hindsight, so one would expect for the features they did choose to implement, that they would be superior in both design and implementation. Sadly, neither are the case except for a few minor exceptions.
> Btrfs doesn't make a distinction between a Clone and Snapshot, they are one and the same with a flag.
Yep, and this is one design choice which on the face of it is straightfoward and convenient, but has the side effect of being very inefficient. Because ZFS snapshots are owned by the dataset, AFAIK there's little refcounting overhead; you're just moving blocks to deadlists based on simple transaction ID number comparisons. If you modify a block and its transaction ID is greater than the latest snapshot, you can dispose of it, otherwise you add it to the snapshot deadlist (and also add the new updated block). If you delete a snapshot, you do the same thing: for each block, if the block transaction ID is later than the transaction ID of the previous snapshot, you dispose of it, else you move it to the previous snapshot's deadlist. No refcounting changes except to decrement for disposal. You only start paying the overhead when you create a clone. This makes ZFS snapshots very cheap, and clones a bit more expensive. Btrfs is always expensive as far as I understand.
Your particular uses might not take advantage of this, but it's something to bear in mind.
> Tell me more how I can mix disks of differing size in RAID on ZFS
You can have pools with vdevs of different sizes (I have one right here). It doesn't make sense to have different sizes within a vdev.
The need for cobbling together different sized discs appears to mainly be something needed for tinkering and testing. No one is going to care about this for production systems. It's a neat feature which few people care about in practice. I'd rather they had spent the time on making the basic featureset reliable.
> > some operations can be performed like snapshotting, but not deletion > See user_subvol_rm_allowed mount option, available since Kernel 3.0
Nice to see some option for this. It's better than nothing, but it's not really equivalent. ZFS has a fine-grained permissions delegation system which is inherited through dataset relationships, rather than coarse capabilities.
> And BTRFS doesn't require it's own separate cache that is apart from the normal filesystem cache.
Not a particular concern for me; it's well integrated on FreeBSD, and it's not a problem in practice on Linux nowadays IME. Do you have a specific problem with the ARC?
Luckily this is only Redhat, not Btrfs itself.
I believe that 2015 estimates from the IDC[1] had RHEL at ~60%, SLE at ~20%, Oracle Linux at ~12% and "Other" at ~8%. But I can't access the document at the moment.
There are not that many companies doing the same. Which is why they have a lot of influence over the direction of things.
Btrfs is definitely not gone from Fedora, for example.
(Disclaimer: I'm on the virtualization team at Red Hat).
I'm still using Btrfs on my backup system, but that's only because I like the dedup enough to overlook the brief hangs.
For those with morbid curiosity on the many stability issues with btrfs as a container file system, this is chronicled in Github: https://github.com/concourse/concourse/issues/1045
I recently setup a software mirroring raid with btrfs and I'm loving features like checksumming. It makes me feel my data is quite safe and can't bit rot anymore. So far it is working fine.
Did you notice that the official description of RAID-1 is "Mostly working"? Are you aware if one of your drive fails, you have one chance to re-mirror it, before the remaining drive can no longer be mounted read-write and you need to dump the filesystem and re-create from scratch?
It might be an inconvenient restriction, but when a drive fails I'm already happy that there will be no data loss.
There's a reason my server is running ZFS on FreeBSD. I also love jails, which let me have as many virtual servers as I want without any virtualization overhead.
It's documented on the status Wiki. RAID 1 is "mostly working". If a mirror drops to having one disk, you can mount it once as a read-write volume (required for resilvering); after that, you have to trash it and start again.
I also had issues with a vms using xfs too.
But, I do use xfs on ssd raids(in our servers, being used for testing) and never had an issue there.
Given our use-case, multi-tenant containers, there weren't many choices which had sane snapshotting support, as well as quotas and some level of subtrees. ZFS on Linux has its own share of issues. I won't say that Btrfs was without issues -- there is still a lot of performance work that needs to be done, especially in a multi-tenant workload, but it looked like there were solutions available.
XFS is an excellent filesystem, and it may work well for our usecase in the near future. It's exciting to see new XFS features landing, like reflinks, and collapsing ranges. Hopefully, folks like Redhat continue trying to bring XFS to the future.
The next fs to make a difference, post ZFS, is likely HAMMER2[0] - it's supposedly already stable for single node use (the ZFS / XFS / ext4 use case), and is advancing towards the multinode-at-the-underlying-fs-level, a first.
[0] https://gitweb.dragonflybsd.org/dragonfly.git/blob_plain/HEA...
Generally though sadly btrfs just isn't getting the man hours from any commercial sponsor it seems :(
And to me this is just another voice which is septical of the current state of Btrfs. And while Redhat is not the only voice, it is an important voice.
Red Hat is an important voice, but please remember that supporting something as part of an enterprise distribution requires that you have engineers that work upstream constantly on said project. You can't just passively support something.
A while ago, most of Red Hat's btrfs developers moved to Facebook and clearly they decided that it wasn't worth the money to hire more people to support btrfs on RHEL. If they didn't see customer demand for it, why should they burden their kernel team with supporting something that nobody is asking them for? SUSE supports btrfs (and not just as a technical preview) and they can switch to SLE if they really want btrfs.
Just because something isn't shipping in RHEL doesn't mean that Red Hat decided that btrfs was bad. They likely decided that either their customers are better suited with other options, or they don't think the cost of getting more engineering talent would be worth it. Btrfs is still shipping in Fedora.
[I work for SUSE.]
I'm not a fan of systemd, but have really loathed CentOS/RHEL compared to Debian/* for years.
Hopefully a lot because, like it or not, RHEL puts a ton of work into the linux kernel and other userland linux tools. We need them to stick around.
If I'd left out my preference, probably would have had fewer downvotes.
I’m not trying to start a flame war here (honest), but among most “usual web” ops folks I know, the opposite of what you’re suggesting seems to hold and Debian/Ubuntu are looked at as an odd choice. It’s probably hire #1 going with what they know, more than anything, and it could very well be selection bias regarding the people I know. I’d love to see stats.
I’ve been working with CentOS or OEL in my own roles for several years now, and I’ve never chosen it. (Not saying I wouldn’t, just that it’s been there when I get there.)
I could be wrong, but I picked that up years ago somewhere.
But if you look at the usage, north america is generally more redhat-family and europe is generally more debian-family, even in the SuSE patch that is Germany...
A full reinstall is required to move from CentOS to RHEL. Oracle has a procedure to directly convert CentOS, RHEL, and Scientific Linux into a supported Oracle Linux system.
KSplice is the most common reason why this conversion is required and mandated. If downtime can no longer be tolerated for upgrades to openssl/glibc/vmlinuz, this is the only available path.
I've never met an individual (not corporate server) or end-user (desktop) running RHEL though.
Back to topic, I've been following Btrfs... At one point it seemed destined to become the default for most Linux systems, but I'm not sure where it is headed now.
Revenues in FY2017 were almost $3 billion per year.
So I did a bit of digging and found out that CentOS ran on Hyper-V just fine. I installed it and had no real reason to complain since. The selection of packages is rather spartan compared to Debian, but there is a third-party repository called EPEL that makes up for that.
Some things are different from Debian, of course, but nothing big, and the system has given me no problems whatsoever on the performance and reliability fronts.
Security support for a wide range of packages is also a reason to prefer Debian over Ubuntu, since most of the Debian-inherited packages ("universe") are excluded in Ubuntu.
I'm not familiar with the Fedora process, but they seem to have a security team and a system of security advisories, which EPEL does not appear to have. Doesn't sound like the same at all.
Sure, most of the time packages in EPEL (and in Ubuntu universe section) will eventually get security updates, but there is no promise or organization of timely security updates.
How can they do this if a packaging (and especially backporting) a fix requires deeper knowledge of the package which probably only the maintainer has?
https://github.com/g0tmi1k/debian-ssh > There was an #ifndef PURIFY there for a reason. It's because the openssl authors knew that line would cause trouble in a memory debuger like Purify or Valgrind.
Where a debian maintainer screwed the RNG of OpenSSL to make valgrind happy. This made any key generated on a debian or ubuntu system from 2006 to 2008 very easily breakable.
Downstream should never touch packages beyond backporting fixes made by upstream.
Here's another example of upstream vs downstream conflict in debian :
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=477454
Or PHP developers being fed up with both RedHat and Debian messing with their runtime on whims :
https://derickrethans.nl/distributions-please-dont-cripple-p...
This is why I heavily support the desire for a new packaging system targeted at developers: snaps, flatpak. The downside of having multiple copies of the same libraries pale in comparison to giving back power to upstream. Distro maintainers are routinely modifying codebases they don't understand. Allow us to have a standard installation process that can install packages.. made by the developers themselves, upstream. Just like all other operating systems do.
And Debian, unlike RHEL/CentOS, packages a lot more than they can even reasonably maintain. The vast majority of packages in a Debian stable are insecure, the security team simply cannot handle the large amount of software outside of the truly core stuff (kernel, web servers) :
https://statuscode.ch/2016/02/distribution-packages-consider...
If you aren't supposed to use the packaged wordpress, phpmyadmin or node, why is debian distributing those packages? Debian by distributing these things in their repo encourages the naive first time linux user to install them through their facilities.
If you're not in a tech business, and are therefore using other people's software, then you want to be on the same platform that they were during development which would have been years ago. If it's Linux, then it's something like RHEL.
I'll ask the opposite question how many people in Enterprise don't use RHEL/CentOS?
> I'm not a fan of systemd, but have really loathed CentOS/RHEL compared to Debian/* for years.
Those two are orthogonal. A lot of people loathe Windows and Oracle yet they are raking in billions every year in licensing.
Then there is Government as well, try selling them a production running in Ubuntu or Arch. It can be done but it won't be very easy. There RHEL is king as well.
Although other systems were running Ubuntu, we decided to go with CentOS 6 for these systems - mainly because it was preferred by the current era of sysadmins. It also had packages for Ruby 1.8 or at least the correct dependencies (I can't remember exactly). And even better, although it was released in 2011, it is still supported until 2020.
If I was going to do the same today I'd just use Docker, but it was fairly new back then.
A majority of Telcos, Banking, Military, Stock Exchanges, Medical and large commercial enterprises
Why?
Because when a system starts having issues at 4am dealing with x amount of transactions per second and the shit is going to hit the fan, they want top class support on hand and not to be awaiting on someone replying on a irc channel / mailing list.
I personally am an Arch user on my home machine and work laptop, yet servers that run commercial workloads and have SLAs tied to them will always run RHEL for the reasons above.
Of course for private usage it's a little ove rthe top. But RH also offers solutions there. There's an upstream OS with all the cool stuff: Fedora. And a downstream OS that is just as stable as RHEL but free: CentOS (not sure if I get the name right. There are so many CabcOS out there nowadays).
"Done for the long-term" is a major overstatement IMO.
If I wait another decade, will btrfs have matured? Will there be ANY half-modern filesystem for linux? I'm not convinced.
Currently bcachefs seems more appealing but well, long way to go there as well.
Sorry, but I don't believe enough in btrfs to try it out for real (and don't have time to play with it just for fun). Especially when playing with more advanced features, the status page does not inspire confidence.
The paragraph on btrfs in https://www.patreon.com/bcachefs seems spot on and exactly the feeling you get after spending a decade of hope on btrfs. And that kind of review is exactly what you don't want on the brand new finally-we-can-store-data-properly-on-linux solution.
btrfs, which was supposed to be Linux's next generation COW filesystem - Linux's answer to zfs. Unfortunately, too much code was written too quickly without focusing on getting the core design correct first, and now it has too many design mistakes baked into the on disk format and an enormous, messy codebase - bigger that xfs. It's taken far too long to stabilize as well - poisoning the well for future filesystems because too many people were burned on btrfs, repeatedly (e.g. Fedora's tried to switch to btrfs multiple times and had to switch at the last minute, and server vendors who years ago hoped to one day roll out btrfs are now quietly migrating to xfs instead).
If I was in ops I would not use btrfs on anything but lab experiments... if i can get hardware RAID, then I can blame the vendor and the equipment
Where did you hear this? This is news to me. https://linuxcontainers.org/ and https://github.com/lxc/lxc both seem active and supported by Canonical to me. The only deprecated project by them is listed to be CGManager.
To RedHat: do something for that menu at the left. It stays in the way when scrolling and zooming. The X to close it is not immediately visible on a small screen. Expected behavior: the menu scrolls away with the page and doesn't stay fixed in the way of the reader.
To Mozilla: open the page with Opera and copy what you see there. Tldr: autofit the text in the screen width. Maybe Chrome does the same. Btw, reader mode doesn't kick in.
It's utterly terribly designed to be an open source information page.
Do you remember their names? Thanks.
Nope
A real shame, especially for those coming from Presto-based Opera Mobile, which used to do reflow seamlessly and in a better layout than modern Blink-based Opera.
Fortunately, Firefox supports addons and you'll find at least a couple ones claiming to do the trick: https://addons.mozilla.org/en-US/android/addon/fit-text-to-w... https://addons.mozilla.org/en-US/android/addon/text-reflow/