At the same time, a few person startup should probably focus on what allows them to deliver the fastest. I’ve seen that work with relatively monolithic systems and with SOA, tooling choice makes a huge impact.
At the same time, a few person startup should probably focus on what allows them to deliver the fastest. I’ve seen that work with relatively monolithic systems and with SOA, tooling choice makes a huge impact.
We had a disk array go sideways, which is when we learned that some dumb motherfucker had put our wiki on the same SAN with production traffic. You know, the wiki where you keep all your run books for solving production issues? Everyone was furious and that team lost some prestige that day. How dumb do you have to be?
Maintenance is looking at all of the probability < 10^-4 issues that are just waiting for you to roll the dice enough times to eventually lose to the birthday problem. Every day you’re lowering the odds that tomorrow will be the day everything burns, because doing nothing is just a waiting game.
This experience ended up being the beginning of the end for the anti-cloud element at the company. Which is too bad because I like having people who understand the physics of our architecture. Saves me from doing all sorts of stupid things myself.
Pretty much all of our docs, and everything concerning more than one team is stored in an instance of our own software running on the software plattform. And it works well.
However, the core operational teams document their core operational knowledge in different git-based systems, all of which are backed up into the archiving as well. This way, if we really lose access to all documentation, we probably lost all workstations of a team to some incident, as well as a repository host, as well as two archive hosts in different parts of europe. At that point, the disaster recovery plan is a bar, to be honest.