AWS IAM is having issues again
twitter.com
twitter.com
They can rebuild that whole datacenter a hundred times over and still not feel it in their wallet.
Agreed.
And usually it's old things, like really really old if something is available in us-east-1 and not us-east-2. I would imagine some ancient ec2 instance types.
Are you saying us-east-2 would be better than us-east-1?
Although IAM service is global one and it is located in us-east-1, so you wouldn't be able to avoid this particular issue.
e.g. we set up our L@E functions in us-east-1, run our actual systems in eu-west-1, but 95%+ of our traffic comes in via eu-west-2, and the L@E functions output logs for that traffic to the eu-west-2 CloudWatch.
For Cloudfront stuff, it's not very region dependent.
I used to work for another pretty big company and we knew to roll out changes to the smaller DCs first.
If their goal is to reduce the number of affected locations (and not care about number of people) then deploying to us-east-1 makes more sense.
Blah.
Maybe, possibly, hopefully...
UPDATE: Status page now shows it https://status.aws.amazon.com/#
https://gaslighting.me is the non-editorialized version.
I've done a lot of cloud site HA work on large sites.
Waiting out the cloud outage so far ends up being the best solution for almost all companies, from both engineering and business standpoints. Eat the outage, but continue with a known-working site afterwards. You just blame the cloud provider for the downtime.
> I've heard of larger companies running a completely redundant hot standby on another independent cloud platform and switching DNS over to the standby when something goes wrong.
In theory, that makes sense. It practise, it almost never works.
If by "independent cloud platform" you mean another AZ or region in the same cloud, that is often attempted and can work reasonably well. If you mean failover from AWS to GCP, then that's unlikely, since everything is different.
An example is whenever DynDNS goes down for 2-3 hours, and everybody builds out flaky failover tools that are less reliable than their original DNS provider - and have to be maintained forever. Might work with one or two domains, becomes a huge ongoing problem with dozens of them. Also, DNS mgmt. APIs are flaky in several dimensions (availability, versioning, parameters, etc.)
Another is that you can't failover to another location that doesn't have all your data, current certificates, monitoring, etc., and the failover site needs the capacity of the original site to work. That costs ongoing money and time, and you never know how well the failover will work or how it will perform.
An example is Heartland, one of the biggest US payment providers, who failed over to another location and took a 5 day outage. Or gitlab, who took a one day outage because of database isues.
I have (automatically) failed over a large site for a publicly-traded company from one AWS region to another, but that took a year of work to setup, and I understood almost all aspects of the site. Afterwards, I realized that almost nobody really has time to organize that either at a conceptual or engineering level. And organizations don't recognize Herculean efforts like that, so think twice beforehand.
Key point: always involve your DBA from the beginning when doing a project like this.
The other day we started using Access Advisor, and we found some of our KMS key policies with a Principal of '*'.
It wasn't marked as globally open, so we planned to fix them a little later.
This morning we found that status had changed.
While we were in the wrong to begin with, it was a little surprising to find the interpretation of the key policy changing overnight.
Of course it became our top priority and is now fixed. Something to look out for...
[0] - https://aws.amazon.com/about-aws/global-infrastructure/regio...
GovCloud I believe is technically separated in this regard, I don't think IAM credentials can operate across partitions.
Though it does seem to be possible to link the existence of a standard AWS account with GovCloud, the same is not true of the China partition
Has it's own version of everything, physically segmented, even for global systems such as IAM.
You have to be a US Citizen to work on it and you need special security clearances.
GovCloud customers only need be a US person or entity, beyond that any further regulatory alignment is up to the customer. AWS does not audit the IAM user base for nationality or any compliance requirements.
Disclaimer: I am an AWS Public Sector Solutions Architect.
So it's entirely possible to depend on IAM but not in a way which this is breaking.
Please confirm and I will work to spread the news - can I cite you as the source?
But genuinely AzureAD is really good. Authentication and Authorization systems are difficult, AD has always been begrudgingly the most well rounded (yes, it has warts) but AzureAD really is nice.
I especially like that my local admin can delegate Enterprise Apps to me so I can create SAML/OAUTH2 SSO links between stuff we use without needing the keys to the kingdom.
I recently set up Enterprise Federation with GCP and it took less than a day. (compared with many months in my last company which used on-prem AD)
https://docs.microsoft.com/en-us/windows/win32/secauthn/cred...
Caveat, I've been out of the Windows/AD game for 7+ years!
(significantly higher == 100% availability, it hasn't gone down... yet)
Similar issue comes with estimating how reliable things are. People are more likely to respond "I had an issue with X too, here's my story" rather than "all good, nothing to report".
When (not if) it does go down, however long it takes to replace is downtime that has to be amortized over the entire X years, to get an accurate picture. If it takes > 260 minutes to detect, replace, and fully recover, then you've already blown the four nines budget you "saved up" for the past five years, when amortizing.
Side-note from a recent twitter thread: If you're not rigorously monitoring availability (preferably with canaries/probes), it's probably much lower than you think. (Using the royal 'you' here, not pointing any fingers.)
Now I host everything on managed services. It was the right decision then, I was broke. Now, though, the cost of having to deal with it is higher.