To get ahead of the cynics — it would not serve the least generous of Google's objectives to be less than transparent about downtime — people figuring it out while the dashboard is green looks much worse.
A lot of customers judge their perception of the severeness of an issue based on the service dashboard. That's why the dashboard makes everything look more rosy than reality.
Every time it says "some customers may be experiencing elevated latency", it usually means "all customers are seeing all requests timeout, and the only ones who aren't are the ones not using the service right now".
Yes seriously - when they say "<5% of users are having trouble accessing Gmail" they are calculating that based on the percentage of all Gmail accounts which have seen an error. Obviously the vast majority of accounts are inactive at any one time, so aren't seeing anything...
Also SLAs and service credits are tied to these official notices which causes even more delay before status updates are approved.
I think I'd prefer a status page that reflected that, or even no status page at all with an explanation, rather than a lying one, though.
The on-caller for the specific service decides on how to produce.
The status dashboard is something which will be (manually, yes by hand) updated by an operations employee, who is a couple of layers behind the actual SWE/SRE who gets the page.
It's complicated, and the human factor is massive in an unpredicted scenario.
I'm an IM myself, and regularly have to make the call regarding status updates, and also seen how other companies do it.
One of my favourites was a big it infra company sending a notification about an outage 6 hours after the event, after we opened and resolved a ticket with them. Was actually pretty useful but made me laugh all the same
It is frustrating that these pages are not updated in real-time, but I do understand wanting to be sure before publishing a message to your entire userbase that you are experiencing disruptions.
Gee, you mean we just discovered that area where Google admits that either automation or machine learning does not actually work?!
Google doesn't need an internal party to hold responsible for the concerns of free mail users or their own contractors who are called "content creators." In part because Gmail users and YouTubers don't lead fortune 50 companies and play golf with Google Execs like Google Cloud customers do.
"Sometimes, it's probably almost approximately around what we think could be along the spectrum of known values." - Me, as CEO.
You need to draw a line in the sand somewhere, but whatever measure you choose is going to be somewhat arbitrary so I think it's reasonable to have a human make the final call (based on some known criteria).
Then put that on the dashboard! It certainly is an outage for the affected users (by your scenario).
- 999/1000 Data centers up
- 100,000 Users without service
- 6 paying customers without service (will get money back)https://mobile.twitter.com/uhoelzle/status/13093135569956618...
So, "kind of" confirmed, I think. Assuming a bad update is the most likely culprit for a whole pool of servers to all crash at the same time.
I've suggested to our team they copy/paste locally anything they're currently editing.