Our datacenter UPS was tied right into the 3-phase output and didn't recognize that the power was still flowing from the city. The diesel engines fired up and started powering the datacenter (redundantly). When the gas ran out a few hours later, everything started shutting down. It was terrible. Hundreds of servers with 600-800+ days uptime all went down and stayed down for hours. We were shuttling people from the office down to the datacenter to fsck and bring machines back up. It was a long night and long next week issuing credits to customers who had lost data.
Today I've got servers with Verio, Bluehost, and LayeredTech. They've all been great for the most part, but none of them have been near perfect. I know that most hosting companies I've worked for or dealt with put uptime near the top of their list of priorities because they know it's an easily measured number people bank on. I'm sure RS will get their act together and be as solid as they've typically been in the past.
Do HA plans not cover this contingency? Do they not specify that someone immediately calls a fuel truck the SECOND the generator kicks on?
It was truly an unexpected case (which was quickly fixed in the UPS hardware), but you know, life is one unexpected case after another, often coming in waves and in bizarre combinations. I don't believe in 100% uptime anymore :)
The word you are looking for is incompetence.
An UPS, even a datacenter scale UPS, is not exactly rocket surgery. If your story is true and the device failed over to diesel without notifying anyone then that's not only an epic engineering failure but also an epic fiscal failure for whoever is liable (perhaps the UPS vendor).
Damages from a full DC blackout easily run into the hundred thousands of dollars per hour, not even counting the unbillable shockwave of "our website is down" multiplied by hundreds of customers.
It's a ridiculously expensive "Oops" that easily dwarfs the cost for deploying a proper UPS with proper testing and proper procedures in first place.
Btw the CAT in our Level3 datacenter over here has a big horn and a flashlight on the side. My naive self wants to believe they are there for situations such as the diesel running out...
I also learned to cut people (and some companies) a little slack when I've been down the road they're on. We had a lot of customers who didn't cut us any slack, took their refund and left, which was their prerogative to do so, but found out for themselves that no company has 100% uptime (many came back, lucky for us).
Edit: Actually just checked, and apparently my host switched from the 99.9999% to 100% as well.
Before you buy from a web host, check their SLA and ask how they cover SLA breaches. Don't do business with a company that won't put their money where their claims are.
My host has:
If Liquid Web is or is not directly responsible for
causing the downtime, the customer will receive a
credit for 10 times ( 1,000% ) the actual amount of
downtime. This means that if your server is unreachable
for 1 hour (beyond the 0.0% allowed), you will receive
10 hours of credit.
Rackspace has: Rackspace Guaranty: We will credit your account 5% of
the monthly fee for each 30 minutes of network
downtime, up to 100% of your monthly fee for the
affected server.edited for clarification: slicehost is owned by rackspace. thought it would be worthwhile to mention since a lot of people here use it; i know i double-checked just to make sure.