To get ahead of the cynics — it would not serve the least generous of Google's objectives to be less than transparent about downtime — people figuring it out while the dashboard is green looks much worse.
To get ahead of the cynics — it would not serve the least generous of Google's objectives to be less than transparent about downtime — people figuring it out while the dashboard is green looks much worse.
It's complicated, and the human factor is massive in an unpredicted scenario.
I'm an IM myself, and regularly have to make the call regarding status updates, and also seen how other companies do it.
One of my favourites was a big it infra company sending a notification about an outage 6 hours after the event, after we opened and resolved a ticket with them. Was actually pretty useful but made me laugh all the same
A lot of customers judge their perception of the severeness of an issue based on the service dashboard. That's why the dashboard makes everything look more rosy than reality.
Every time it says "some customers may be experiencing elevated latency", it usually means "all customers are seeing all requests timeout, and the only ones who aren't are the ones not using the service right now".
Yes seriously - when they say "<5% of users are having trouble accessing Gmail" they are calculating that based on the percentage of all Gmail accounts which have seen an error. Obviously the vast majority of accounts are inactive at any one time, so aren't seeing anything...
Also SLAs and service credits are tied to these official notices which causes even more delay before status updates are approved.
I think I'd prefer a status page that reflected that, or even no status page at all with an explanation, rather than a lying one, though.
The on-caller for the specific service decides on how to produce.
The status dashboard is something which will be (manually, yes by hand) updated by an operations employee, who is a couple of layers behind the actual SWE/SRE who gets the page.