Zrepl – ZFS replication
zrepl.github.io
zrepl.github.io
ZFS is much more than just a filesystem and related tools. It also supports backup/restore and disk partitioning and partitioning of data via filesystems (not AFAIK available with EXT4.) And probably other things I either don;t know about or have forgotten to mention.
That said, it's great reading its success stories and praises, and my next system will use ZFS. :)
I have a "ghetto NAS" of 24 drives plugged into my server via USB, and getting a raidz3 set up on there was one command:
`zpool create tank raidz3 drive1 drive2 drive3 ....`
Doing scrubs and replacing disks are pretty easy, just using `scrub` and `replace` commands with zpool.
I haven't really felt the need to leave ZFS, at least not in regards to RAIDs; on root I still use ext4, though I might change to btrfs on my next install.
Storage Spaces with ReFS has had the most impressive feature set from this perspective. Add and remove an arbitrary number of drives of arbitrary sizes and it'll use the full capacity of the drives to whatever parity level you set it to. It has its own downsides of course, on top of being Windows only, but it's the only FS/pool combo that has really made me think "ZFS doesn't have quite everything perfect".
There's some current work to add a disk to existing vdevs, and I think even a semi-working PR for it now: https://github.com/openzfs/zfs/pull/15022. Hopefully in theory that will make ZFS a little less frustrating.
That's my main issue with ZFS. It does too many things, and is too clever/magic for my taste. Many things can go wrong when relying on a monolithic system with that much complexity. I much prefer the Unix "do one thing well" approach, and mixing purpose-built tools to suit my needs, rather than using one tool for everything.
You're a lot more likely to lose/corrupt data with ext4 than ZFS. Ext4 will happily corrupt data silently. The core conceit of ZFS is it doesn't trust the underlying hardware. ZFS even allows duplicated data on a single disk. You lose capacity but gain robustness.
When ZFS support in Linux was still new, I wouldn't dare rely on it, and it's the same reason I avoid btrfs today, even though its features are appealing. But now that ZFS is quite stable, and after hearing its praises for years, I do want to give it a try. :)
You might not know. How are you validating the integrity of every file regularly?
In the 90s I lost some files to corruption which went unnoticed for years and propagated to every backup, so by the time I went looking for the file I had years worth of backups of those files, all corrupt. This is one of the reasons zfs is such a happy place.
Using a smart filesystem would be an improvement, but it comes at the expense of less flexibility and more complexity, and, until recently on Linux at least, relying on unstable software.
Where the hell are you getting this from? What instability has existed in ZFS on Linux? Are you confusing it with BTRFS or something? I think your fears of ZFS are seriously misplaced.
lol. Nothing is stopping you from doing manual backups using ZFS. One should never rely on just one backup anyways, if the data is critical. For me, snapshots are a great way to protect from "oh, I accidentally deleted this folder", which ext4 doesn't have. Yes, you can use replication to sync these snapshots somewhere else, but nothing is stopping you from continuing to do manual backups on a file level. It's just a file system, after all. So it doesn't really make sense as a justification as to why you are hesitant to use ZFS. In fact, it's one of the reasons I liked it so much: While you can do all these cool things with it, you don't have to. It doesn't pressure you to use these features. If you're ready, they're there, but until then, it's just a file system, and a very robust one at that.
If you have an irrational fear of data loss (it's really quite rational), zfs is the only place you should be comfortable.
I've been on zfs since its earliest days (~2004-ish, inside Sun at the time) and never lost a single byte ever since. Computers and hard drives have died but zfs just rocks on.
I like the ability to incrementally send only the changes of an encrypted filesystem to a target server that never had the encryption key at anytime.
I'm using this to share backup space with my friends. We both push/pull our encrypted snapshot diffs every hour. I don't have my friends' keys so I can't read their data, and they can't read mine. In case of emergency, I can go to their place with my key, and recover my data from their systems.
* dd to copy the boot sector (to a ZFS filesystem) * rsync to copy the boot (FAT32) and root (EXT4) partitions (also to a ZFS filesystem) * syncoid to copy the ZFS filesystem to another host.
This process runs daily. When a system craps, it's pretty straightforward to restore the boot sector and boot and root partitions and finally the ZFS pool.
I'm curious zrepl does that might be useful for me.
I'm only replicating the data filesystems.
From what I see, sanoid looks quite similar to zrepl. Both tools are probably able to achieve similar results.
I do feel that zrepl has more features so far, though. But hey, if your setup is working and your data is secure, that's the most important.
From a user perspective, I'm still feeling a lack of UX in restoration/clone-management from the snapshots synced by zrepl - things can easily get messy if you don't know exactly what you're doing and keep track of things properly. That might be out of scope for zrepl itself, though.
Some more proactive/pre-emptive pruning would also be great - falling behind on syncing can be downright painful if you don't pay attention... Which is just to say, don't skip setting up some basic monitoring with alerts for production loads (:
Either way, all things considered zrepl is the nicest zfs replication option I found so far.
https://www.psc.edu/hpn-ssh-home/hpn-ssh-faq/ has a patchset that fixes the problem.
When I choose libraries or tools for work, let's say I need a Go library for something; I use the star count and the commit activity together to determine whether a project is popular enough (vs others) to use over the others.
I don't look at stars or fork count. To see if I should use a tool or a library, I'll be seeking technical blog posts and reviews.
I don't like how software development is centering around GitHub, and how it is handled like a social network.
Wha's more, you can't know why someone added a star. Maybe they found something random cool in the project, or found it looks cool but haven't really tried, and most importantly, you don't know if they added a star after some time, building some solid experience with the library or the tool.
This metric doesn't even exist for projects not developed on GitHub and many worthy ones are indeed elsewhere. And I suppose many of us don't use stars on GitHub, so you are only ever going to measure "popularity" aka "I liked something here" of only a part of the population that is willing to star things.
It's also probably a self-sustaining process: people are not going to give stars to alternative projects they haven't tried... because the first one had more stars in the first place! This feature is probably quite misleading in the end.
In a previous job, the number of forks and stars, as well as recent commits played an important choice when choosing a library. I hated this. It lead to some bad choices and technical dept.
In short, when you are considering stars, you are looking at a self-sustained process, sustained only by people who participate in it, and that is GitHub-specific (excluding interesting projects that happen to not be hosted there). This thing probably pushes the actual popularity of some projects for reasons that are likely very random, at times. GitHub is not making us a favor with this feature.
Commit history or activity in the bug tracker is somewhat important (though sometimes a project doesn't receive so many commits because it is stable enough). The social media gamification things on GitHub like counts and badges should just be ignored, don't fall for it and its addictive aspect.
I’ll probably give zrepl a try based solely on that.
I like this script but never heard from it and probably will not look at it for a while.
Only issue i've found is the fact there is no official Debian repository and i always do forget to recompile it after upgrading Debian :-D
I really like the configuration system integrated in ZFS set/get...
https://news.ycombinator.com/item?id=32014969
Here's a nice summary someone wrote:
Also had a dead 6TB hard disk this morning: WD60EZAZ. Tried replacing the PCB as I had a few functional ones (ex RAID array), happened to have the right security bit, but no dice. You can hear it try to take-off repeatedly but it never achieves altitude. It's been awhile!
So you can make some fancy things like (relatively) cost-free snapshots, and some extra data safety with per-file checksumming. It also uses WAL (similarly to databases) which allows (similiar to databases) replication to remote system.