It appeals to my limited knowledge and non-existant experience that this would be a solution to the prevention of this occurring again in the future?
It appeals to my limited knowledge and non-existant experience that this would be a solution to the prevention of this occurring again in the future?
My 2cents:
Any change should be considered dangerous, and be tested first, time weighted to it's level of change. (An internal policy that could be communicated publicly. One that I apply to all my staff)
It would be also good to have data center clusters (preferably a datacenter clusters are sharded among regions) which would allow these changes to happen as necessary. A random cluster being the "first cluster" with a roll back in place if fails, or proceed through to other clusters progressively until all are live.
The sharding should hopefully alleviate any corner of the world taking any massive hit due to degraded performance.
If they did some colos as vendor J and some colos as vendor C, I think it would be manageable, but I don't really know how much of the cross colo traffic is actually their routers talking to their routers. Homogeneity in networks makes things easier to manage, until a platform fault breaks everything at the same time. In this case, at least it was related to a change they had made and happened quickly, so it was easy to determine the cause; other platform faults may not be as easy to determine, but if only your vendor X colos fell over, at least you'd have your vendor C colos up and something to look for.