Was it Metra or the CTA out of curiosity?
why is McDonalds dependent on AWS for their app to work? (/s)
Why shouldn’t Government IT be using the same tools regular companies use for IT?
A service that is unable to handle such a failure does not qualify as being ready for deployment. And I'm not just saying that to be a self-righteous pedant, I'm saying it because this kind of failure is statistically likely to happen at some point, so ignoring it is setting yourself up for major (and potentially very expensive) problems if you don't test for it regularly.
One major outage per year doesn't necessarily mean just a few hours of downtime, it could mean having to redeploy your entire service somewhere else, which could take several days or more if you haven't prepared for it. How many of those services can live with that?
Apparently, many can! Look at how many companies choose not to pay ransom when hit with a ransomware attack, or prefer to negotiate for days instead of buckling straight away, even if it means operations are crippled for weeks. They don't typically go bankrupt afterwards, everyone coped, life moves on.
And even more now: http://techblog.netflix.com/2013/12/active-active-for-multi-...
The companies we talk here about are most likely running a monolith app on Centos6..
But an entire region? I've never worked anywhere that decided being multi-region was a good tradeoff. At best we've replicated data to another region and had some of our management services there, so in the absolute worst case (which would need to be much worse than yesterday) we could rebuild our product's infrastructure there.
Do I agree with this approach? Not in all possible cases of course, but for my employers? Overall, yes. It mitigates the highest impact risks. Going further would have significant complexity and costs. Those companies success or failure haven't been impacted by their multi-region strategy AFAICT.
Services that truly need 100% uptime (and by that I don't mean what management say they WANT, but they are prepared to PAY) are a tiny minority, to the point that I imagine most software developers never work on one.
Even setups that i have seen to handle such cases, usually had some single point of failure somewhere.
And lest face it. Even if you do multi region on AWS, that won't protect you next time someone screws up BGP/switches/DNS or whatever and every region goes down for a while. In that case you better have failover to some other cloud vendor. Even planes or other safety critical equipment is not 100% failure proof.
If you're going to concentrate risk on AWS, it better be essentially flawless, stable, and highly-redundant.
> All service interfaces, without exception, must be designed from the ground up to be externalizable. That is to say, the team must plan and design to be able to expose the interface to developers in the outside world. No exceptions.
1: https://nordicapis.com/the-bezos-api-mandate-amazons-manifes...