gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
In my company we are split in 3, US, EU, APAC, and we have the same issue with global outage for stuff we could have just managed regionally. For all the savings of the global architecture, they disappear each minute a client is down on a global outage because a guy thousands of kms away messed up.
You dont have to unify, at all. You dont unify with your competitors, and the world has not exploded: compete internally between regions ?
As far as global services go though, it's easy enough to say "it should just not be possible", but how do you propose doing that in practice for a global service?
How does new config going to go out, globally, without being global? How do global services work if they're not global? How does DDoS protection work if you don't do it globally?
People make fun of "webscale" but operating Google is really difficult and complicated!
The global outage thing seems to be a consistent "feature" of GCP - how are we supposed to architect our deployments if the regional isolation model is not a bulwark against high availability on GCP?
As I understand it, GCP is already designed to make global outages impossible. Obviously this outage shows that they messed up somehow and some global point of failure still remains. Looking forward to the post-mortem.
How much more work would Google create for themselves if they had not globalized their stack? Are we talking something like 5 subsets to manage instead of 1?
Ex-googler, no particular knowledge of this event, information might be out of date.
1. Some networking-related service has global, non-standard (compared to the rest of the company) configuration
2. The relevant VP is aware and has decided not to change it because that change is quoted as impossible
3. Some change elsewhere happens that assumes standard configuration
4. The networking service breaks and causes a global outage
5. VP is told to fix it
6. Fix rolls out in weeks, because it wasn't as hard as they said before
These constraints get thrown to the wind when the downtime is already happening.
Of course, if you deploy a change to all of your separated stacks at once through some sort of automated pipeline it doesn’t matter too much. Easy to break everything simultaneously that way if there’s some difference between test and prod you didn’t realize was there.
I reckon the only to achieve that would be to have the same level of interoperability between regions as you would get between two distinct cloud providers.
Of course, at Google scale 'partial' is still very big.
And why enterprises clamoring for AWS to feature match Google's global stuff (theoretically making I.T. easier) instead of remaining regionally isolated (making I.T. actually more resilient, without extra work if I.T. operators can figure out infra-as-code patterns) should STFU and learn themselves some Terraform, Pulumi, or etc.
Also, AWS, if you're in this thread, stop with the recent cross-region coupling features already. Google's doing it wrong, explain that, and be patient, the market share will come back to you when they run out of the GCP subsidy dollars.
> gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
If that's what you really need, then distribute your assets across GCP, AWS, and DO. That likely means not using any cloud-specific features such as Lambda. AWS is actually really good in this regard, as SES and RDS are easily copied to regular instances in other cloud providers, that possibly wrap some cloud-specific feature themselves.