Let’s Talk OpenZFS Snapshots
klarasystems.com
klarasystems.com
I sometimes blow away all snapshots if I have a big clear out or rewrite of data, or else that stuff sticks around for months.
Fun fact, the ZFS send-receive functionality was born with snapshots being created every few minutes in mind:
> ZFS send and receive wasn’t considered until late in the development cycle. The idea came to me in 2005. ZFS was nearing integration, and I was spending a few months working in Sun’s new office in Beijing, China. The network link between Beijing and Menlo Park was low-bandwidth and high-latency, and our NFS-based source code manager was painful to use. I needed a way to quickly ship incremental changes to a workspace across the Pacific. A POSIX-based utility (like rsync) would at best have to traverse all the files and directories to find the few that were modified since a specific date, and at worst it would compare the files on each side, incurring many high-latency round trips. I realized that the block pointers in ZFS already have all the information we need: the birth time allows us to quickly and precisely find the blocks that are changed since a given snapshot. It was easiest to implement ZFS send at the DMU layer, just below the ZPL. This allows the semantically-important changes to be transferred exactly, without any special code to handle features like NFSv4 style ACLs, case-insensitivity, and extended attributes. Storage-specific settings, like compression and RAID type, can be different on the sending and receiving sides. What began as a workaround for a crappy network link has become one of the pillars of ZFS, and the foundation of several remote replication products, including the one at Delphix.
* https://www.delphix.com/blog/delphix-engineering/zfs-10-year...
That is, the remote destination does not need to have the key/passphrase to the file system to have an (encrypted) copy of the data.
If your production server goes down, you can restore the encrypted file system without the encryption key, and only when you try to mount the restored ZFS file system will you be prompted for the key/passphrase.
> This means that you can use ZFS replication to back up your data to an untrusted location, without concerns about your private data being read. With raw send, your data is replicated without ever being decrypted—and without the backup target ever being able to decrypt it at all. This means you can replicate your offsite backups to a friend's house or at a commercial service like rsync.net or zfs.rent without compromising your privacy, even if the service (or friend) is itself compromised.
* https://arstechnica.com/gadgets/2021/06/a-quick-start-guide-...
* https://www.rsync.net/products/zfsintro.html
* https://www.rsync.net/resources/howto/snapshots.html
Whether their ZFS implementation supports OpenZFS encryption is something you'll have to e-mail them about. They have an HN account:
* https://news.ycombinator.com/user?id=rsync
Edit: Yes, it seems that they do:
Yes, we do.
If you get a zpool from us (which is a bit different than a "normal" rsync.net account) it is running the latest stable ZoL codebase and, thus, supports encryption and "raw send", etc.
However ...
As elegant and performant as 'zfs send' is, the benefits are realized in large datasets where efficiency and performance really matters. If you're just sending 200 or 400 or 800 GB of data to an untrusted destination (like rsync.net) you should probably just use borg[1][2].
In order to give you a zpool of your own we need to give you a full blown VM (bhyve) with resource guarantees and your own IP address, etc.
So there is a 1TB minimum and no discounts.
Alternatively, if you just get a plain old rsync.net account the minimum account size is much smaller and if you're an expert and don't need any (borg specific) support there is a discounted plan[1].
If OpenZFS supported some kind of "virtual zpool" feature (I guess it doesn't exist yet) that you could provision out of your main ZFS infrastructure and hand off to the customer, would that be useful to reduce the resource overhead related to offering ZFS directly?
Do you have a sense of why dedicated/isolated infra (VM & IP) might be necessary to provide this service?
We strongly encourage people to use ipv6, which we support everywhere ...
If I remember correctly, the rsync protocol is actually very nearly optimal in terms of minimising round trips.
(For anyone wondering, https://github.com/bahamas10/zfs-prune-snapshots)
The whole article does not touch the subject of boot environments ...
Maybe the author can update the page and state that the article is about managing snapshots instead of boot environments?
ZFS boot environments are awesome. You can get them on both FreeBSD and Illumos with beadm. On Linux there are a number of utilities, but I've only used ZSys. It comes working out of the box if you do a ZFS-on-root install with Ubuntu.
Fyi: snapshots are read-only. Normal, read-writable filesystem can be created by `zfs clone root@snap newroot`.
2021-07-29-10-hour
2021-07-29-10-53-minute
or similar. Then it's just a matter of ensuring that you have NTP working and your clocks are in sync.Can you give more details so I know what to watch out for? (Recently started running Docker+ZFS in prod, haven't seen it break yet but appreciate heads up)
I have to run snapd on one of my servers and I hate everything about it.
Am I in for a world of hurt?
1) updates are pushed/applied and services restarted on an arbitrary schedule which means you can't test updates prior to rolling them out or otherwise control when a service is taken offline by an update
2) snapd doesn't log to standard OS logging facilities and the logs it does have are cleared after 72 hours
3) snapd pollutes home directories with a directory called "snap"
snapd is a poorly conceived project written by people that are profoundly ignorant of how real services and servers are actually operated. Despite having the above flaws repeatedly pointed out, the developers have shown zero interest in taking actual concrete actions to address them.
How would you work around it in a production setting if you HAD to use it?
Auto-updating sounds like the biggest issue of those.
More worryingly, 'prune' locked up for about 10 minutes before getting to do it's thing. pinning the cpu the whole time. This was on my devbox, so maybe it's my doing but did not inspire confidence.
I've run FreeBSD ZFS for about a decade and openzfs/ubuntu in prod for kafka clusters so I'm no stranger - this docker usage was just particularly strange.
It might actually have been this ms sql issue that led me to do that, actually:
My question is on snapshots vs file-specific backups: If a folder is set up for non-destructive writes, does it then need snapshots? Is there a benefit to using both snapshots and NDR in the field?
By using ZFS Boot Environments.
With Boot Environments (I don't know if Linux based Distros let one do this) on Solaris and Illumos, we have multiple active root file systems to switch between and can mark snapshots for each of them.
Certainly, for Belenix (KDE based opensolaris distro), we snapshotted the entire root file system.
FreeBSD stole this feature from Solaris as well:
IMHO implemented or ported is better word here.
bieaz: https://openzfs.github.io/openzfs-docs/Getting%20Started/Arc... zbm: https://github.com/zbm-dev/zfsbootmenu zsys: https://github.com/ubuntu/zsys
Sanoid makes snapshots according to a schedule you define (e.g. 1 monthly, 31 daily, 24 hourly). It also prunes old snapshots to ensure that you always have the exact snapshots you need available and no more.
Syncoid uses zfs send/recv to copy over snapshots from one ZFS dataset to another, typically between machines. You basically have two options: either you can:
1) copy over ALL of the snapshots from the source in the order they were created (like cherry picking one git commit after another onto another branch), or you can:
2) configure it to make a new snapshot every time it syncs and just copy over the diff since the last time you synced (a bit like squashing all the commits into one and just cherry-picking that).
The former is good if you want to maintain the same snapshots on the source and the destination (you can also use sanoid on the destination to prune the snapshots further).
The latter is good if you want to minimize the amount of churn sent over the network. For example if you're snapshotting VM images you'd be interested in sending over the state of the VM as it appears at the end of the day, rather than all the intermediate state.
- take snapshots hourly and daily
- prune old snapshots (--keep option)
- send the snapshot to a remote host (--post-snapshot=zfs send ...)