100 Lines of C, in a Closet (2020)
blog.notryan.com
blog.notryan.com
If one wants to avoid the downsides of just checking in big binary blobs, git-annex[1] is fantastic for maintaining archives (not backups). It also tracks locations of files (which of my drives is this file on?) in a completely peer-to-peer eventually consistent manner. It has a nearly unreasonable amount of features in multiple layers, though.
(Tip: set annex.largefiles[2] to avoid messing up your repo when you inevitably type `git add` instead of `git annex add`.)
I just used Cloudflare, and a Go package [1] to lock down access to anyone but Cloudflare (because at the time their tunnels required Argo).
I now also use a VPS that's dirt cheap [2] if you make it IPv6 only (I have 3 in 3 different availability zones for about 1.5€/mo, it's insane). Cloudflare exposes it as IPv4+IPv6.
Compute light stuff comes from the VPS, compute heavy is proxied home instead.
[1]: https://pkg.go.dev/github.com/ncruces/go-cloudflare/origin
Yes! It can take hours or days if you're moving a large machine with tons of traffic, too. Everything is fast for small n.
This doesn't make it less cool, though.
Also, we do these migrations not because of HA failovers. If we need that kind of resilience, we'd either setup an active/standby or active/active HA VMs in the first place.
Instead we do this because machine maintenance and/or upgrades. A RAM stick dies, but the server tolerates it, so we leisurely drain the server for maintenance, or we replace the nodes with more powerful ones, so we decommission it by draining, take it out of the cluster, add the new one and remigrate stuff back in.
Also, most of the VMs we migrate doesn't know or have the capability of multi-instance execution. They are designed to work as solo machines. This is another reason we prefer to migrate them.
Lastly that's a multi-tenant system. We're not the admins of many of the systems sitting on that particular cluster, yet we proudly hit >99.98% uptime every month.
TL;DR: Migration is a different modus operandi and is not for scaling or active HA 99.9% of the time. We move these machines around to service the hardware running these, and nobody notices anything.
I remember moving a huge NFS partition which was experiencing constant writes to another server with DRBD. With almost zero downtime. So is a nice tool if you want to move a filesystem with such huge amount of files that even iterating over the file tree takes hours
I’d like to be able to create zpools for DBs to take advantage of snapshotting, but combining ZFS and Ceph on the same underlying disk (even if they are enterprise NVMe drives) is fraught with peril.
https://github.com/MarquisdeGeek/ultra
Maybe it'll inspire someone!