Unless there are 2x a year restore tests, I personally assume a 60% backup fail rate.
Unless there are 2x a year restore tests, I personally assume a 60% backup fail rate.
Downtime happens to pretty much every service out there. In this case the company was incredibly forthright and so we can make fun of their stupid mistakes, but really most mistakes are stupid when you look at them -- when you make thousands of decisions a day, some of them will seem silly in hindsight. We just never learn about most of them.
It's difficult and expensive to build and maintain a solid system, and even if you want to, time and financial pressures often just don't let you, on top of the issue of just communicating the need for solid engineering, as it usually only becomes apparent when the problems start occurring.
I had a customer once in the 90s take a 30 hour outage that cost them nearly $6M in fines because some asshole put a budget freeze on anything related to cleaning, including tape drive cleaner carts. The dopey ops guy kept using one tape on multiple drives, making them do nothing.
I could personally rattle off a dozen stories like this at late stage startups, Fortune 10 and .gov.
The only reason many businesses are alive is luck and reliable SAN.
However, the interview seems to suggest that "write everything down" is the way to solve this problem and was what allowed for all their growth and success. So the solid process they have built is to write everything down, which they obviously didn't do or they wouldn't have had a 7 layer disaster recovery failure.
I agree... but I guess I read that coming from the perspective of someone who isn't a super fan of distributed teams. Many teams think that futzing around in slack or hangouts is enough. "Writing things down" is practically discovering the wheel to some places!
First, no one has the discipline to write much of anything down.
It's a very boring process and it's always going to be low-resolution. Your most meticulous documentation writers will get fired for failing to get their "real work" done. If you personally recognize the value of the documentation they furnish and thus refuse to fire them, all of their peers will feel that they are dead weight, which is still a ticket out of the company.
Second, after all that work, no has the discipline to read much of anything that you've written!
They skim. They glance. They Ctrl+F. They don't read. When you're in a pickle and you can pick out a life-saving bit of documentation, it's amazing, but that happens quite rarely and requires a lot of energy.
How many times have you pulled up docs, tried to follow them, gotten a really confusing error that you spent hours trying to troubleshoot only to find out that there's a one-sentence explanation tucked away in the third sentence of the fourth paragraph on the page you originally pulled up? This just happened to me _last week_, and frequently the most frustrating problems are small things like that.
People don't read. It's nothing personal, they just don't read. It takes a lot of cognitive energy. People are biologically programmed to conserve as much as energy as possible. Good programmers are both lazy and dumb!
If you want documentation that means something, it needs to be part of the process of actually working. I don't mean you need to add "write docs" to your checklist, I mean meeting the operational standards should be the only way things can get done in the first place.
The operating procedure needs to be married to the actual completion of the task, and that means setting good baseline project standards and setting reliable enforcement on those standards.
Code should be self-documenting to a reasonable extent. Tests should be mandatory. Peer reviews and signoffs should be mandatory. Internal company discussions should be recorded and referenceable. Documentation only works when it's self-generating.
In short, it should be run like a mature open-source project with an open IRC channel, mailing list, bug tracker, commit history, mandatory tests, maintainer signoffs, merge processes, code and style standards, docstrings and good automated documentation generators, and so on.
There is an easy solution to dealing with people who refuse to read things. (Hint: It doesn't involve recording hours of meetings.) You need standard operating procedures to document what people do at an appropriate level of detail. Period.
Operational procedure is different than code documentation.
The goal is to marry the SOP and the actual completion of the task, to block off the shortcuts and require a minimum expenditure from external discipline/motivation reserves to get people to follow those SOPs.
This is why it's baked into the company culture. It's very common for someone to ask where something is in the handbook or if an issue has been created for something.
> It's a very boring process and it's always going to be low-resolution. Your most meticulous documentation writers will get fired for failing to get their "real work" done. If you personally recognize the value of the documentation they furnish and thus refuse to fire them, all of their peers will feel that they are dead weight, which is still a ticket out of the company.
Fwiw, everyone is responsible for maintaining the handbook/our process and procedure documentation. The docs team isn't on the hook for it, nor is it the sole responsibility of engineering.
> Second, after all that work, no has the discipline to read much of anything that you've written!
After enough reminders, you'd be amazed at how quickly people learn to RTFM at work.
> They skim. They glance. They Ctrl+F. They don't read. When you're in a pickle and you can pick out a life-saving bit of documentation, it's amazing, but that happens quite rarely and requires a lot of energy.
It's true that people skim the handbook (the guide with all of the "Here's how you get access to Twitter accounts" stuff), but runbooks are actually looked at when stuff hits the fan. I think it's important to differentiate these two things since they have different purposes. Imo runbooks should be as lean as possible for that very reason.
> Good programmers are both lazy and dumb!
This level of documentation is helpful for non-developers who make up a significant part of many organizations. We're not just talking about documenting code here.
> If you want documentation that means something, it needs to be part of the process of actually working.
100%, this is the only way it works.
I find that in practice, with non-mission-critical activities, especially documentation, when everyone is responsible for getting it done (And making it useful), no-one is.
7 layers of backup processes, that they had written down, all failed and went unchecked, by any of their employees across the 160 locations.