More on today's Gmail issue
gmailblog.blogspot.com
gmailblog.blogspot.com
I like his little veiled pitch for Google's services when he talks about how easy it was to bring more request routers online given their elastic architecture. It makes me wonder why that elasticity isn't automated -- more routers should automatically be brought online if any routers hit their maximum load.
The correct solution, as outlined in their post, is for the servers to slow down when overloaded instead of trying to push load onto another server.
Now imagine we had a magic "start more servers" mode. We push new builds once a month, say. We'd just keep spinning up servers automatically without anyone noticing probably. Soon we'd hit some other system bottleneck (I'd bet something like too much DB traffic), and then we'd be off on a wild goose chase wondering what was wrong with the database or something.
(These are semi-hypothetical. As a team we discussed whether we wanted to have our monitoring processes spin up new frontends and decided that we should probably keep humans in the loop ;))
Compare to:
http://developer.amazonwebservices.com/connect/message.jspa?...
http://en.wikipedia.org/wiki/Northeast_Blackout_of_1965
"it's all just a little bit of history repeating"
Also, the outage, for me anyway, seemed to last much longer then the stated 100 minutes. I seem to remember being unable to access GMail for a span of about 3 hours today.
I'm super impressed that they responded so quickly. When a first notification comes in, it is rarely with a fully fleshed out diagnosis. It is usually some runner/monitoring tool sending you an alert saying something is busted. If you get lucky, you had enough monitoring tools across the system to atleast get close to the source. If not, you have some debugging ahead of you. And then, you need to figure out what a good fix is and given that the entire system is in a wobbly state, a bad fix could make the situation worse. And then, you figure out how to rollout the fix and actually make the change. If you're smart, you'll do it in a staged manner and be able to roll back the moment something goes wrong.
In short, these things take time. Going from notification to working fix in 90 minutes for what sounds like a nasty network hardware issue is very good.
[um, I work for Yahoo, but do you really need a disclaimer on a compliment?]
I work for a web company, and a 100 minute outage is considered absolutely outrageous, but of course it happens.
But that's the spin, and I suspect you're querying the preference. And you're right - whether something was down for 100 minutes or 1 Hour and 40 minutes wouldn't make much difference to my life.
You know who got blasted out of sleep at 3 AM because my email was down? I don't know, but it sure as heck wasn't me.
I did too, and I didn't get much sleep in those days.
Only people who were unfortunate enough to be working on the GMail team from UTC+7, perhaps UTC+6, instead of in Mountain View.
In other words, this type of outage is pretty unlikely to ever plague Google again.
"Well, if my email is down, then my competitor's is too!"
I don't really know how much your argument applies to companies looking for hosted email; if they can't send, they can't send, so it doesn't matter if someone else can't receive. In fact, probably makes it better b/c they look less individually incompetent.
GE's central email farm went down when I interned there a couple summers ago. All 300k+ employees' emails went out.
The business is still affected by the outage if they host locally, there will probably just be more of them. And there's the illusion of control: GE had noone to blame but themselves (and the incompetent contractor who hit the big red button), whereas they could've pinned this on Google and felt worried about external risks and dependencies.