Yeah, I'm not worried about being targeted in an RCA and pointedly asked why I chose a region with way better uptime than `us-tirefire-1`.
What _is_ worth considering is whether your more carefully considered region will perform better during an actual outage where some critical AWS resource goes down in Virginia, taking my region with it anyway.
AWS Organizations/Account management is us-east-1.
And if you want a CDN with a custom hostname and want TLS…you have to use us-east-1.
CloudFront CDN has a similar setup. The SSL certificate and key have to be hosted in us-east-1 for control plane operations but once deployed, the public data plane is globally or regionally dispersed. There is no auto failover for the cert dependency yet. The SLA is only three 9s. Also depends on Route53.
The elephant in the room for hyperscalers is the potential for rogue employees or a cyber attack on a control plane. Considering the high stakes and economic criticality of these platforms, both are inevitable and both have likely already happened.
There's just not much motivation left to do better systems.
So if you tried to be "smart" and set up in Ohio you got crushed by the thundering herd coming out of Virginia and then bit again because aws barely cares about you region and neither does anyone else.
The truth is Amazon doesn't have any real backup for Virginia. They don't have the capacity anywhere else and the whole geographic distribution scheme is a chimera.
Makes one wonder, does us-west-2 have the capacity to take on this surge?
“Duh, because there’s an AZ in us-east-1 where you can’t configure EBS volumes for attachment to fargate launch type ECS tasks, of course. Everybody knows that…”
:p
However: Don’t underestimate community support (in the areas you’re likely to want it) when comparing development stacks.
What’s the point in having 64 Gb of DDR5 and 16 cores @ 4.2 GHz if not to be able to have a couple electron apps sitting at idle yet somehow still using the equivalent computational resources of the most powerful supercomputer on earth in the mid 1990s.
Oh and put everything behind the strictest cloudflare settings you can, so that even a whiff of anything that’s not a Windows 11 laptop or iPhone on a major U.S. network residential or mobile IP gets non-stop bot checks!
Is this from real experience of something that actually happened, or just imagined?
The only things that matter in a decision are:
* Services that are available in the region
* (if relevant and critical) Latency to other services
* SLAs for the region
Everything else is irrelevant.
If you think AWS is so bad that their SLAs are not trustworthy, that's a different problem to solve.
Separately from that, if you are trying to move certain types of non-mainstream IBM workloads to cloud (AIX, IBM i, z/OS) then IBM is tier 1 in that case
us-east-2 is objectively a better region to pick if you want US east, yet you feel safer picking use1 because “I’m safer making a worse decision that everyone understands is worse, as long as everyone else does it as well.”
If you never get blamed for a US east outage, that's better than us-east-2 if that could get you blamed 0.5% of the time when it goes down and us1 isn't down or etc
I can’t tell if it’s you thinking this way, or if your company is setup to incentivize this. But either way, I think it’s suboptimal.
That’s not about “risk profile” of the business or making the right decision for the customer, that’s about risk profile of saving your own tail in the organizational gamesmanship sense. Which is a shame, tbh. For both the customer and for people making tech decisions.
I fully appreciate that some companies may encourage this behavior, and we all need a job so we have to work somewhere, but this type of thinking objectively leads to worse technology decisions and I hope I never have to work for a company that encourages this.
Edit: addressing blame when things go wrong… don’t you think it would be a better story to tell your boss that you did the right thing for the customer, rather than “I did this because everyone else does it, even though most of us agree it’s worse for the customer in general”. I would assume I’d get more blame for the 2nd decision than the 1st.
See any companies getting credit for it in the last AWS outage? I didn't. My employers didn't reward vendors who stayed up during it.
Shame about your employer, though.
US-East-2 staying up isn’t my responsibility. If I need my own failover, I’m going to select a different region anyway.
And it’s not like US-East-2 isn’t already huge and growing. It’s effectively becoming another US-East-1.
No, but you can be blamed if other things are up and yours is not. If everyone's stuff is down, it is just a natural disaster.
If my cloud provider goes down and also takes down Spotify, Snapchat, Venmo, Reddit, and a ton of other major services that my customers and my boss use daily, they will be much more understanding that there is a third party issue that we can more or less wait out.
Every provider has outages. US-east-2 will sometimes go down. If I'm not going to make a system that can fail over from one provider to another (which is a lot of work and can be expensive, and really won't be actively used often), it might be better to just use the popular one and go with the group.
The regions provide the same functionality, so I see genuinely no downside or additional work to picking the 2 regions over the 2 regions.
It seems like one of those no brainer decisions to me. I take pride in being up when everyone else is down. 5 9s or bust, baby!