I sometimes blow away all snapshots if I have a big clear out or rewrite of data, or else that stuff sticks around for months.
I sometimes blow away all snapshots if I have a big clear out or rewrite of data, or else that stuff sticks around for months.
Fun fact, the ZFS send-receive functionality was born with snapshots being created every few minutes in mind:
> ZFS send and receive wasn’t considered until late in the development cycle. The idea came to me in 2005. ZFS was nearing integration, and I was spending a few months working in Sun’s new office in Beijing, China. The network link between Beijing and Menlo Park was low-bandwidth and high-latency, and our NFS-based source code manager was painful to use. I needed a way to quickly ship incremental changes to a workspace across the Pacific. A POSIX-based utility (like rsync) would at best have to traverse all the files and directories to find the few that were modified since a specific date, and at worst it would compare the files on each side, incurring many high-latency round trips. I realized that the block pointers in ZFS already have all the information we need: the birth time allows us to quickly and precisely find the blocks that are changed since a given snapshot. It was easiest to implement ZFS send at the DMU layer, just below the ZPL. This allows the semantically-important changes to be transferred exactly, without any special code to handle features like NFSv4 style ACLs, case-insensitivity, and extended attributes. Storage-specific settings, like compression and RAID type, can be different on the sending and receiving sides. What began as a workaround for a crappy network link has become one of the pillars of ZFS, and the foundation of several remote replication products, including the one at Delphix.
* https://www.delphix.com/blog/delphix-engineering/zfs-10-year...
That is, the remote destination does not need to have the key/passphrase to the file system to have an (encrypted) copy of the data.
If your production server goes down, you can restore the encrypted file system without the encryption key, and only when you try to mount the restored ZFS file system will you be prompted for the key/passphrase.
> This means that you can use ZFS replication to back up your data to an untrusted location, without concerns about your private data being read. With raw send, your data is replicated without ever being decrypted—and without the backup target ever being able to decrypt it at all. This means you can replicate your offsite backups to a friend's house or at a commercial service like rsync.net or zfs.rent without compromising your privacy, even if the service (or friend) is itself compromised.
* https://arstechnica.com/gadgets/2021/06/a-quick-start-guide-...
* https://www.rsync.net/products/zfsintro.html
* https://www.rsync.net/resources/howto/snapshots.html
Whether their ZFS implementation supports OpenZFS encryption is something you'll have to e-mail them about. They have an HN account:
* https://news.ycombinator.com/user?id=rsync
Edit: Yes, it seems that they do:
Yes, we do.
If you get a zpool from us (which is a bit different than a "normal" rsync.net account) it is running the latest stable ZoL codebase and, thus, supports encryption and "raw send", etc.
However ...
As elegant and performant as 'zfs send' is, the benefits are realized in large datasets where efficiency and performance really matters. If you're just sending 200 or 400 or 800 GB of data to an untrusted destination (like rsync.net) you should probably just use borg[1][2].
In order to give you a zpool of your own we need to give you a full blown VM (bhyve) with resource guarantees and your own IP address, etc.
So there is a 1TB minimum and no discounts.
Alternatively, if you just get a plain old rsync.net account the minimum account size is much smaller and if you're an expert and don't need any (borg specific) support there is a discounted plan[1].
If OpenZFS supported some kind of "virtual zpool" feature (I guess it doesn't exist yet) that you could provision out of your main ZFS infrastructure and hand off to the customer, would that be useful to reduce the resource overhead related to offering ZFS directly?
Do you have a sense of why dedicated/isolated infra (VM & IP) might be necessary to provide this service?
We strongly encourage people to use ipv6, which we support everywhere ...
If I remember correctly, the rsync protocol is actually very nearly optimal in terms of minimising round trips.
(For anyone wondering, https://github.com/bahamas10/zfs-prune-snapshots)