The Meltdown That Brought Our Startup to Its Knees for 15 Hours
groovehq.com
groovehq.com
Also, only one master database? No dumps to offsite storage? No secondaries? Crazy talk. This is SysAdmin 101 stuff.
> No longer will infrastructure be a “feature” to be > weighed and prioritized against others in our backlog. > It’s the foundation of everything we have, everything > we do, and it will be treated as such."
I'm glad someone is learning this lesson. I wish it wasn't under those circumstances.
One of the things our server monitoring system does is text and phone (thanks Twilo!) when things go this far south. Of course in Alex's case it might not have helped since his phone was dead and not charged but at least one of the team would have gotten the message. The only down side for me is that when my family on the east coast texts me something in the "morning" which is like 4AM pacific, I bolt awake thinking its a server outage.