"DBA" is a full-time position, not an addon to a developer's duties. They are separate skillsets; you don't get a "2-for-1" special by hiring an expert developer + DBA in one person for one lowly salary.
Do you realize this event would have never happened if you had hired a pure DBA? Or do you really believe you can pin the blame on your ruby developers for not being able to wrangle a production postgres database?
The oblivious or intentionally cheap "the DBA must be an amazing ruby developer" expectation shows your hiring staff - or the management guiding them - has absolutely no clue what they are doing. I can just imagine the internal discussion right now; pointing the finger at the developers with no postgres experience, or downplaying the significance of this event and pretending like it was simply bad luck, and lying to yourselves about how "it will never happen again".
This job posting is completely outside the realm of reason. If that job posting has been up for months or years, I can see its description being exactly why you didn't have the right talent on board to avoid this incident.
This is a mistake. The number of DBAs who are good with postgres is very small compared to something like mysql. The truly talented pool for such a position is too small to expect them to also be a developer. The very mention of terms like "ruby" and "programming" should be removed from that job post. It's not a realistic expectation.
> I'm the CEO of GitLab https://about.gitlab.com/ More information about me is on http://sytse.com
https://about.gitlab.com/2015/04/08/the-remote-manifesto/
But maybe you meant part-time?
It is always a HUGE red flag when a company opts to have all or a majority of their workforce working remotely. It's a cost-cutting measure, nothing more. Cutting costs equates to cutting corners, and the business - and its customers - suffer the deserved consequences.
This really explains the flippant "it's 11pm and I want to go to bed" reaction in their report. The guy doesn't have an office to go to when shit hits the fan. He's sitting at home, with a bloody ssh terminal open, trying to remotely debug critical engineering problems over a slow vpn connection. Alarmingly huge red warning flags.
Might wanna revise that today.
Instead of backing up every system in isolation, have well though out backup/restore processes for all parts of their operation.
Saying that as someone who's done exactly this before (professionally, for mission critical places). ;)
eg:
· Inventory the systems (boxes, services, etc)
· Determine what each needs (package dependencies, etc)
· Create scripting (etc) for consistent backups
· Work out the restore processes
· Make it work (can take several test/dev iterations)
· Document it
And also (importantly):
· Have the ops staff perform the documented processes, to reveal holes in the docs, and show up parts which need simplifying