The problem here was that there seemed to be no backup and restore strategy at all.
That said, though: if you're a tiny group like this and don't have the time to invest in a gold plated storage design, "Just Back Up All The Servers" is a pretty reasonable choice to make.
in that role I would typically back up the deploy directory (/opt/foo) and the /etc/ and parts of /var/ every day to Glacier.
Of course this is all pre-docker and pre-k8s but this is 2013 we are talking about
Yes, if you have a deployment strategy where you can reproducibly deploy from your source code repository, then you don't need to back up the servers. But until you've tested that that works, and I mean really tested it, you need server backups.
In case of things like this. Also it's usually quicker to recover a server using an image than it would be to provision from source in the case of, say, a hardware failure.
What would you do when you had 100 servers?
Something more scalable. There's no reason why you should pick an approach to solving a problem and stick with it forever. If the situation changes you do something more appropriate. That doesn't mean the small scale approach is wrong when the scale is small.
Because the rule number one of the programmers club, is backup everything, everywhere, all the time, if you don't remember do it again, and if you're sure you have enough backups do it again anyway. I even do backup of my backups everyday (in pendrives, in cds, in ancient scrolls, etc)
There's a certain elegance and assurance you get from this that has been lost with the times, akin to how monolithic server software with all functionality natively available in the code has gone away in favor of microservices. Now you have message queues, k/v stores, caches, search engines as a microservices that are tacked on to the core services and rarely fully understood by the engineering team and containing more functionality than the codebase ever really utilizes. Ends up being more complicated in manage in a lot of ways. I think the emergence of microservices is one of the driving forces behind selective state backups, because you can never back up the entire state at once, everything is too spread out. You're not going to back up the running state of the k8s node, or whatever
Keep in mind that once the servers are treated as cattle then individual server backups are typically no longer needed. The developer in this case would likely not be in this situation then.
As one writer said early in the unraveling of Agile, it's like you were told that if you eat your vegetables you can have desert, and now we just want desert but don't want to eat our vegetables.
While working on other things I discovered modifications written directly to cached versions of the Django instance itself, which were not committed to any source control. He simply edited the instance of Django in the disk image, which was then replicated to our various deployments.
The worst part was the mods this guy made were perfectly capable of running within our project, from code in source control. Why he took the approach he did was a complete mystery to us.
However, the support staff still had root logins to all our servers, so they followed the letter of the law by not having their patches go through the code review / source control / gradual rollout process...
And, "Do you have a QA team," he said, expecting the answer "no".
God, the many times I had told them so months before difficult situations arrived, just to be swiftly ignored by business because some devs didn't want to bother, and wanted to get to more shiny and fancy and interesting problems.
Nor rolling it and both testing this procedure and putting people off mutating the server.