Random aside: any chance you are related to the Calculus on Manifolds Spivak?
He says that any good theorem is worth generalizing, and I've generalized that to any life rule.
looks like only four 9's
> looks like only four 9's
That's why the Germans are such good engineers. Did the drives fail? Nein.
Did the CPU overheat? Nein.
Did the power get cut? Nein.
Did the network go down? Nein.
That's "four neins" right there. 5N = 99.999%
3N = 99.9%
1N5 = 95%
5N is <43m12s downtime per month.Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.
Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it).
I'm sure there is some subtlety to this, but it does mean that large corps with influence should be talking to AWS to ensure that status information corresponds with actual service outages.
I think you mean "start thinking they can get what they pay for"
2. Cost is not an issue (until it is but you’re already locked in so oh well)
3. Faang has drained the talent pool of people who know how
If this needed a CEO to eventually get around to pressing a button that said "show users the actual information about a problem" that reflects poorly on amazon.
let me repeat that: my AWS trainign that is run by AWS that I pay AWS for isn't working, because AWS is having control plane (or other) issues. This is several hours after the initial incident. We're doing training in us-west-2, but the identity service and other components run in us-east-1.
Partial availability isn’t the same as no availability.
i'd say when other companies who run their infrastruture on AWS are going out, it's hard to argue it's not a real outage.
But AWS status _has_ changed to yellow at this point. Probably heroku could be completely down because of an AWS problem, and AWS status would still not show red. But at least yellow tells us there's a problem, the distinction between yellow and red probably only matters at this point to lawyers arguing about the AWS SLA, the rest of us know yellow means "problems", red will never be seen, and green means "maybe problems anyway".
I believe the entire us-east-1 could be entirely missing, and they'd still only put a yellow not a red on status page. After all, the other regions are all fine, right?
edit: partly down; it's sporadically failing
Versus, if a company has X SLA contracts signed, that point to Y reimbursement for being out for Z minutes, so it's easily calculable.
Mom: Alexa, did you break something?
Alexa: No.
M: Really? What's this? 500 Internal server error
A: ok maybe management console is down
M: Anything else?
A: ...
A: ... ok maybe cloudwatch logs
M: Ah hah. What else?
A: That's it, I swear!
M: 503 ClientError
A: ...well okay secretsmanager might be busted too...
Me: Alexa, is AWS down right now?
Alexa: I'd rather not answer that
That's a bit like involving your kid in an argument between parents.
I don't think a region being down is something that you can be unsure about.
Hopefully it results in a class action lawsuit for enough money that Amazon decides that an automated system is better than trying to supply human judgement.
This is what Amazon, the startup, understood.
Step 1: Always make it right and make the customer happy, even if it hurts in $.
Step 2: If you find you're losing too much money over a particular issue, fix the issue.
Amazon, one of the world's largest companies, seems to have forgotten that the risk of not reporting accurately isn't money, but breaking the feedback chain. Once you start gaming metrics, no leaders know what's really important to work on internally, because no leaders know what the actual issues are. It's late Soviet Union in a nutshell. If everyone is gaming the system at all levels, then eventually the ability to objectively execute decreases, because effort is misallocated due to misunderstanding.
How come an action of a private company in a capitalist country is like the Soviet Union?
It’s very frustrating. Why even have them?
Scratch that... console is not loading at all now :)
Amazing and scary to see all the unrelated services down right now.
Step 1: deploy status checks to an external cloud.
That being said, AWS status pages are up.
When the biggest cloud provider in the world is famous for gaslighting, it sets expectations for our whole industry.
It's fucking disgraceful that they tolerate such a lack of integrity in their organization.
My favorite status page, though, is Slack's. You can read an article in the New York Times about how Slack was down for most of a day, and the status page is just like "some percentage of users experienced minor connectivity issues". "Some percentage" is code for "100%" and "minor" is code for "total". Good try.
So what's a good health check actually report these days? Is it just about its own status, or should it include a breakdown of the status of external dependencies as part of its folded up status?
"We are experiencing an error. Our apologies – We will be back up soon."
It is very unlikely that Amazon would deliberately make your messages cross the Atlantic just to find an American region that is unable to serve you.
We can't even update our product to say it's down, because accessing the product requires a process that is currently dead.
> AWS Management Console Home page is currently unavailable.
> You can monitor status on the AWS Service Health Dashboard.
"AWS Service Health Dashboard" is a link to status.aws.amazon.com... which is ALL GREEN. So... thanks for the suggestion?
At this point the AWS service health dashboard is kind of famous for always been green isn't it? It's a joke to it's users. Do the folks who work on the relevant AWS internal team(s) know this, and just not have the resources to do anything about it, or what? If it's a harder problem than you'd think for interesting technical reasons, that'd be interesting to hear about.