Zrepl on Rsync.net
blog.lenny.ninja
blog.lenny.ninja
If you want to zfs send to an rsync.net account you need to have an account enabled for that which is the same price, but has a higher (4TB) minimum account size.
It used to be 1TB minimum until pretty recently but they went with 4TB now.
If you want a less space solution, just spin up a t3.small (I suppose at least 2GB memory would be better for zfs) on AWS as Ubuntu with Cold HDD (minimum 125GB) as extra storage which is quite cheap at about 0.015/GB and install "zfsutils-linux" and you're good to go to use zfs.
It's probably fine for a true "last resort" backup, but if you're frequently restoring it's worth keeping in mind.
Cold storage has a complicated pricing model because it's al-a-carte, but offsite backups tend to be write-once-read-never.
$0.0036 per GB / Month for aws $0.007 per GB-month for storage for gcs
Well, if you're being smart, it's going to be write-once-read-sometimes-once because you should be testing your backups every so often. You'd probably want to backup to some kind of service in the middle, then at the end of the month/quarter ship a backup to cold storage.
If you want a generic remote zfs fs it's going to cost $$$. I have that but it's not offsite.
If you don't care about stability or reliability, choose whatever the landing page sells you.
This is random access, live storage - which is not the case with cold storage options such as Glacier.
A more fitting comparison would be S3 ...
I think you are, in many cases, correct - offsite backups can, indeed, be write-once ... but if you're looking for the very specific efficiencies and features that zfs-send affords you, that will not be the case.
It's still fuzzy to me specifically why a dedicated 'vm' is necessary. You don't need a dedicated ssh ip for normal 'shell' accounts, and certainly you still provide security and privacy for these customers. Is it because zfs doesn't have fine-grained permissions to allow doing zfs send/recv actions on one's own datasets/volumes without also giving access to other customers' data? Or is it because send/recv workloads just consume more compute resources in general?
So, basically, you need to be root to fully manage your own zpool ... and if you need to be root, you need to be in a VM. In our case, we use bhyve and we have had very good success.
We have this on very good authority - Allan Jude of Klara Systems has helped us audit the entire setup and there is not, unfortunately, a lighter way to do it.
[1]: https://openzfs.github.io/openzfs-docs/man/8/zfs-allow.8.htm...
The author of 'zrepl' is Christian Schwarz (@problame) - sorry for the mixup :)
I found an old 2017 issue about one user's reasoning when deciding between sanoid and zsnapsend: https://github.com/jimsalterjrs/sanoid/issues/102
And another huge benefit for my setup is the ability to use an HTTPS connection for the replication instead of SSH. My source and target servers are in different continents so there's pretty big latency. zrepl's HTTPS server manages to transfer data at ~890Mbps while SSH doesn't deal with the high latency as well and only manages ~190Mbps.
I tried rsync for this purpose in the past. Our main office gets 150 Mbps down / 20 up. I tried uploading the initial 1TB snapshot and after a week it had not finished. Meanwhile a fair amount of that original snapshot becomes stale during that time.
Are you just supposed to start up these hourly snapshots and hope everything catches up with itself eventually?
Once it was up and running most snapshots took a few minutes to sync, always finished before the morning anyway.
Definite +1 to rsync.net, this was >15 years ago but it was always 100% solid, I don't think I ever had any issues. It's nice to see they're still doing the same thing and haven't bloated it with crap!
The other option, if you're colocating, is to send a seed drive ahead to the datacenter. (I think there was a startup on here a while back where they'd basically colo your drives in their own JBODs, and then charge you a nominal monthly fee for a VPS w/ those drives passed through as a zpool.) You might pay some nominal fee for remote hands, but it beats waiting for terabytes of data to squeeze through your local cableco's wildly asymmetric pipe.