I want cheap and reliable snapshots, export & import of file systems like ZFS datasets, simple compression, caching facilities(like SLOG and ARC) and decent performance.
I want cheap and reliable snapshots, export & import of file systems like ZFS datasets, simple compression, caching facilities(like SLOG and ARC) and decent performance.
Bcachefs has never had an unrecoverable data error AFAIK, even though it isn't even considered stable enough to merge into the kernel. The bcache on disk format won't be considered stable until he merges his code into mainline, though he doesn't feel he will need to adjust it further.
Features that currently work: Full data checksumming Compression Multiple device support Tiering/writeback caching RAID1/RAID10
All of these are stable, tested, and mostly bug free. Honestly, once the code gets mainlined you'll be able to start using it very quickly.
Main issue right now is performance, as it about as slow as BTRFS, which isn't inspiring. However the author has stated that he's going for correctness first, then he'll begin optimizing.
Snapshots are a huge feature for sure, but it's not like bcachefs is completely incapable without them.
There was a very recent update he gave in late December (2019) that mentioned he's actively chipping away at the roadblocks for snapshots.
I have heard of durability issues with btrfs, and do not want to touch it if it fails with its primary job.
Does anyone remember the parity patches they rejected in 2014?
> Your work is very very good, it just doesn’t fit our business case.
I haven't followed it much. Does it have anything more than mirroring (that's stable) these days?
You're saying they should stop supporting a project that was considered stable by the time the other started being developed. Why do that? What makes Bcachefs a better choice?
I don't think too many people consider it stable enough for production, either. (Unless you count a very limited subset of its functionality).
I rather run Bcachefs today than Btrfs, by a mile. At least with bcachefs I won't lose my data.
Even if the data is still on the drive and a bugfix would make the filesystem recoverable again, they don't have the time/knowledge/resources to untangle that codebase and make fixes. Even BTRFS developers don't trust the filesystem with their own data.
If you are on Bcachefs and you encounter an unrecoverable bug, the developer will ask for some logs, or reproduction steps, or potentially even remote debugging access to your corrupt filesystem.
And then he will fix the bug, releasing a new version that can read/repair your filesystem. He knows his codebase like the back of hand.
In my research, I couldn't find any examples of someone actually losing data due to Bcachefs. All the bugs appeared to be "data has been written to drive, but bug prevented reading"
While I would still hesitate to trust Bcachefs, I would trust it way more than BTRFS.
Definitely something to try out (backing up my home servers is just about to reach viability for me, so I'd definitely consider switching to it in that use case).
Thanks!
After that, none of the features like compression, snapshots, COW or checksums meant anything to me. I'm much happier with ext4 and xfs on lvm.
In a way, I would rather it bomb out and declare a total loss than to keep sinking more time into it as it leads you along.
But seeing how so many people had lost data using it, I will never use btrfs...
I think they just said: "The on-disk data structure is stable" and lots of people misinterpreted that as "the whole thing is stable"
A stable on-disk data structure just means it's been frozen and can't be changed in non-backwards compatible ways. It says nothing about code quality, feature completeness or if the frozen data structure was any good.
Or Bcachefs is probably the only thing that might get there.
The amount of engineering hours went into ZFS is insane. It is easy to get a project that has 80% similarity on the surface, but then you spend the same amount of time from 0 - 80% on the last 20% and edge cases. ZFS has been battle tested by many. Rsync is on ZFS.
The amount of Petabyte stored in ZFS safely over the years gives peace of mind.
Speaking of Rsync, normally a topic of ZFS on HN will have him resurface. Hasn't seen any reply from him yet.
I think everyone is in agreement that ZFS can't be included in the mainline kernel. The question is just if users should be able to install and use it themselves or not.
The follow up actually clears things up pretty well. https://www.realworldtech.com/forum/?threadid=189711&curpost...
For import export, IIRC XFS has support for it and you can dump/import LV snapshots to get atomicity.
For caching there is LVM cache, should be again possible to combine with thinpool & RAID. Or you can use it separately for normal LV.
All this is functionality tested by years of production use.
For compression/deduplication, that is AFAIK work in progress upstream based on the open sourced VDO code.
Never made snapshots with LVM. Always used LVM as a way to carve up logical storage from a pool of physical devices but nothing more. I need to RTFM on how snapshotting would work there - could I restore just a few files from an hour ago while letting everything else be as they are?
With ZFS, I use RAM as read chace(ARC) and an Optane disk as sync write cache(SLOG). I wonder if LVM cache would let me do such a thing. Again, a pointer for more manual reading for me.
Compression is a nice to have for me at this moment. Good to know that it is being worked on at the LVM layer.
For some reading about LVM thin provisioning:
http://man7.org/linux/man-pages/man7/lvmthin.7.html
https://access.redhat.com/documentation/en-us/red_hat_enterp...
There is difference between 'all these tools have been used in production' and 'this is an integrated tool that has been used for 15+ years in the biggest storage installations in the world'.
There's also Lustre but it's a different beast altogether for a different scenario.
Once you actually use them, you discover all the ways that btrfs is a pain and zfs is a (minor) joy:
- snapshot management
- online scrub
- data integrity
- disk management
I lost data from perfectly healthy-appearing btrfs systems twice. I've never lost data on maintained zfs systems, and I now trust a lot more data to zfs than I ever have to btrfs.
Granted, at enterprise scale this hardly matters because you can just send-receive to rebuild pools if you have enough spares, but for consumer-grade deployments it's a non-negligible annoyance.
Maybe SV is different...
The indirection tables are survivable for fixing short term mistakes, though.
> I lost data from perfectly healthy-appearing btrfs systems twice.
I still consider btrfs as beta-level software. This is why I never looked into it very seriously and asked this question.
Looks like btrfs has something around five years to be considered serious at the scale where ZFS just starting to warm-up.
Overall:
Device size: 142.86GiB
Device allocated: 48.05GiB
Device unallocated: 94.81GiB
Device missing: 0.00B
Used: 37.75GiB
Free (estimated): 103.94GiB (min: 103.94GiB)
Data ratio: 1.00
Metadata ratio: 1.00
Global reserve: 82.20MiB (used: 0.00B)Twice btrfs ended up in a non-mountable situation, but both times it was due to a known issue and #btrfs on freenode was able to walk me through getting it working again.
With ZFS, I neded up in a non-mountable system, and the response in both #zfs and #zfsonlinux to me posting the error message were, "that sucks, hope you had backups." Since I both had backups and it was my laptop 2000 miles from home that was my only computing device, I didn't dig deeper to see if I could discover the problem. FWIW, I've been using ZFS on that same hardware for almost 2 years since with no issues.
But in the end it always turns out that only if you 'use' it correctly it is actually not gone eat your data.
I used ZFS for far longer and had far fewer issues.
Once a little more guidance comes out about how to properly use VDO and Stratis together, I'll move my personal stuff to it.
There is also BeeGFS, I haven't used it but /r/datahoarders sometimes touts it.
Not for linux but I have been keeping an eye on M Dillons DragonFly BSD where he has been working on HAMMER2, which is very interesting.
I don't know much but bcachefs has been making more waves lately also.
I think the bottom line is that people need to have good backup in place regardless.
Redhat throwing towel on their support for development does not instill confidence either.
Nothing personally against Btrfs. Just an end user making a file system choice saying what I care about.
> People are making a bigger deal of this than it is. Since I left Red Hat in 2012 there hasn't been another engineer to pick up the work, and it is _a lot_ of work.
btrfs still has a write hole for RAID5/6 (the kind I primarily use) [0] and has since at least 2012.
For a filesystem to have a bug leading to dataloss unpatched for over 8 years is just plain unacceptable.
I've also had issues even without RAID, particularly after power outages. Not minor issues but "your filesystem is gone now, sorry" issues.
Pretty much all software-raid systems suffer from it unless they explicitly patch over it via journaling. Hardware raid gets away with it if it has battery backups, if they don't they suffer from exactly the same problem.
https://www.synology.com/en-global/knowledgebase/DSM/help/DS...
Their recovery documentation [0] indicates that SHR is just plain mdadm + LVM and a couple of NAS recovery sites [1,2] indicate the same.
In the end I got a Reddit post [3] with a response from a Synology representative who says that the btrfs filesystem will request a read from a redundant copy from mdadm in order to correct checksum errors.
I wonder whether this is unique to Synology or whether the change has been upstreamed into the main Linux kernel.
[0]: https://www.synology.com/en-global/knowledgebase/DSM/tutoria...
[1]: https://support.reclaime.com/kb/article/8-synology-shr-raid/
[2]: http://www.nas-recovery.com/kb_hybrydraid.php
[3]: https://www.reddit.com/r/DataHoarder/comments/5yb13m/anyone_...
I thought I wanted RAID5, but after reading horror stories of drives failing when replacing a failed drive, I decided it just wasn't worth the risk.
I currently run RAID1, and when I need more space, I'll double my drives and set up RAID10. I don't need most of the features of ZFS, so BTRFS works for me.
If a disk fails and resilvering causes a cascading failure, I can restore from a backup.
I think you might be mistaking RAID for a backup, which is a mistake. RAID is very much not a backup or any kind of substitute for a backup. A backup ensures durability and integrity of your data by providing an independent fallback should your primary storage fail. RAID ensures availability of your data by keeping your storage online when up to N disks fail.
RAID won't protect you from an accidental "rm -Rf /", ransomware or other malware, bugs in your software or many other common causes of data loss.
I might consider RAID10 if I were running a business-critical server where availability was paramount, or where I needed decent random read/write performance but even so I'd still want a hot-failover and a comprehensively tested backup strategy.
Btw, ZFS scrub is not only a RAID-block-check but also a partial fsck, so its not really comparable.
Hardware RAID has very poor longevity. Vendor support and battery backup replacement collide in BIOS and host management badly.
Disclaimer: I work on Dell rackmounts, which means rather than native SAS I am 'Dells hack on SAS' which is a problem and I know its possible to 'downgrade' back to native.
Somewhat recently I dealt with LSI and Dell cards. Longevity seemed just fine for a normal 3 year server lifecycle. The only time we had an issue is when the power went down in the data center. The power spike fried a few of the cards. Luckily we had spares.
Way way back I dealt with the Compaq/hp smartarrays. Those were awful. Also anything consumer grade is awful.