If that's not sufficient, what more are you looking for, and what other large cloud providers consistently meet that standard?
If that's not sufficient, what more are you looking for, and what other large cloud providers consistently meet that standard?
e.g.
https://aws.amazon.com/message/2329B7/
https://aws.amazon.com/message/41926/
To be fair to Google, they haven't had enough time to perform a detailed autopsy, and some GCP incident summaries have shown meat on the bones e.g. https://status.cloud.google.com/incident/compute/16007. And balancing the scales, the AWS status page is notorious for showing green when things are ... not so verdant.
I have seen full <public cloud> internal outage tickets and the volume of detail is unsurprisingly vast, and boiling it down into summaries - both internal and external - without whitewashing, without emotion, to capture an honest and coherent narration of all the relevant events and all the useful forward learnings is an epic task for even a skilled technical writer and/or principal engineer. You don't get to rest just because services are up, some folks at Google will have a sleep deficit this week.
Some examples:
https://status.cloud.google.com/incident/cloud-networking/18...
https://status.cloud.google.com/incident/cloud-pubsub/19001
https://status.cloud.google.com/incident/cloud-networking/18...
https://status.cloud.google.com/incident/cloud-networking/18...
https://status.cloud.google.com/incident/compute/18012
Given that this was a multi-region outage that lasted several hours and impacted a substantial number of services, I'd expect a detailed postmortem to follow.
Half the people in this thread are overlooking that fact and going into outrage mode.
Every time I read a Google post-mortem, they seem to hand wave everything away as "a configuration error", "bug", or "bad deploy" and their resolution always has the generic "implement changes to things" that says absolutely nothing. Honestly, when the the causes of these massive disruptions are so simply dismissed, it portrays their system as frail amateur work.
While it sucks that multiple regions malfunctioned simultaneously for several hours, I can't really fault them for their communication about the issue.