Anyways, you can imagine we had all sorts of attacks and miners or other abusive software running. This on top of using ephemeral nodes for our free service meant hosts were always coming and going and ceph was always busy migrating data around. The wrong combo of nodes dying and bursting traffic and beta versions of Rook meant we ran into a huge number of edge cases. We did some optimization and re-design, but it turned out there just weren't enough folks interested in paying for multi-tenant Kubernetes. We did learn an absolute ton about multi-tenant K8s, so, if anyone is running into those challenges, feel free to hire us :P
Or just a disk with unreliable USB connection?
The worst that comes to mind for me was a node failure in the middle of a major version upgrade. Not likely a big deal for proper deployments, but I don't have enough nodes to tolerate complete node failure for most of my data.
Grabbed a new root/boot SSD, reinstalled the OS, reinstalled the OSDs on each disk, told Ceph what OSD ID each one had previously (not actually sure if that was required), and....voila, they just rejoined the cluster and started serving their data like nothing ever happened.