.... so much for limiting Blast Radius
.... so much for limiting Blast Radius
A good infra architecture has the “blast radius” of any issue confined to only part of your infra fleet. Avoiding the global outage.
Think of it like a navy ship. When a mussel breaches the hull, the ship is designed to contain the leak in a single section. Avoiding the entire ship sinking. Similar pattern is desired in your service and compute infrastructure.
The outage google just faced is equivalent to a single missel taking down an entire navy fleet.
> The outage google just faced is equivalent to a single missel taking down an entire navy fleet.
I really thought you were going to say "is equivalent to a single missile taking down an entire ship" - but your extreme analogy made even more sense in this context.
I knew barnacles were a problem, I hadn't realised there were other dangerous bivalves. :)
As the time of the outage increases, service health probably trends to zero without being able to manage the service... but for a few hours, it's not a disaster. Usually.