‘99.9999%’ uptime – it’s an illusion
hackernoon.com
hackernoon.com
I am a systems admin at a large multi national company. Downtime can run in the billions per hour. I’ve only ever seen billion dollar downtime once in eight years, and it was for a few hours only. The DBAs delayed the failover decision, God knows why. They should have done that immediately.
We have strict change controls and auditing so that the disaster recovery site always acts like it should, just like prod. Replication always occurs and when it fails it gets attention. Failover is regularly tested. Good governance makes all the difference.
This article describes a failure of governance. It is a management issue.
Edit: And governance is free. Replication is now within reach of even super tight budgets. See my reply below. So small companies can do this, which was my intended point.
But replication is everywhere now. MySQL, rsync, whatever. Plenty of choices.
I appreciate the feedback though!
[1] http://www.oracle.com/technetwork/middleware/goldengate/over...
[2] https://docs.oracle.com/cd/B19306_01/server.102/b14239/conce...
The company probably doesn’t have the budget to support a $billion uptime organization with proper “governance”
But I does sound like five 9’s is possible in your case.
Edit: just saw author is claiming “six 9’s” is impossible ... maybe
Also use solid versioning, tested backups, have a tested backout strategy with a secondary remediation path for all changes. Paranoia is cheap :-)
Someone might say, “It’s more work.” And paging engineers at 3am isn’t work? Would you rather burn out your engineers or have them relaxed and focused on responding to a replication issue at 2pm? :-) You know what they say about an ounce of prevention.
“My 10th anniversary of using @DreamHost – hosting over 60 domains with zero downtime thus far. Can I just say it's one of the greatest tools in my life? Their site: bit.ly/2kl6Vr2“
Is it zero downtime? Probably not but my sites haven’t seen an outage that I recall. Amazing support too.
Taking into account common cause failures (which are often as a result of root causes by human factors) your six nines can very easily be more like three or four nines.
As a functional safety engineer that applies IEC61508/61511 for a living this is a common problem and addressed by diversity in systems, achieving the same goal in redundant systems by different hardware/software means.
Airbus and Space Shuttle flight controls do a similar thing with multiple CPUs of different types with software implemented to the same spec by different teams, in some cases different languages.
Not saying this is what they should be doing in their case as the mission is not so critical, but if you really want five or six nines this becomes extremely difficult with limited diversity in hardware and software.
For me there's never been a bug so bad it couldn't wait until monday morning.