The irony or perhaps the tragedy of building a low friction service is that you have to have experts on the lower level high friction stuff.
I would hope that after a couple of hours downtime, they'd bring up a fresh machine with Ansible or whatever. Hardware or AWS/GCP Vm.