AWS us-east-1 outage
status.aws.amazon.com
status.aws.amazon.com
> 8:22 AM PST We are investigating increased error rates for the AWS Management Console.
> 8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US-EAST-1. Customers may be able to access region-specific consoles going to https://console.aws.amazon.com/. So, to access the US-WEST-2 console, try https://us-west-2.console.aws.amazon.com/
aws iam list-policies
An error occurred (503) when calling the ListPolicies operation (reached max retries: 2): Service UnavailableI'm wondering if the cause of the outage has to do with something changing in the way IAM is interpreted ?
STS at least has recently started supporting regional endpoints, but most things involving users, groups, roles, and authentication are completely dependent on us-east-1.
I opened it again just now (maybe 10 minutes later) and it now shows DynamoDB has issues.
If past incidents are anything to go by, it's going to get worse before it gets better. Rube Goldberg machines aren't known for their resilience to internal faults.
IMHO it should mean that the rate of errors is increased but the service is still able to serve a substantial amount of traffic. If the rate of errors is bigger than, let's say, 90% that's not an increased error rate, that's an outage.
> Dev1: Pushing code for branch "master" to "AWS API". > <slackbot> Your deploy finished in 4 minutes > Dev2: I can't react the API in east-1 > Dev1: Works from my computer
I've broken things before and been aware of it, but didn't acknowledge them until I was confident I could fix them. It allows you to maintain an image of expertise to those outside who care about the broken things but aren't savvy to what or why it's broken. Meanwhile you spent hours, days, weeks addressing the issue and suddenly pull a magic solution out of your hat to look like someone impossible to replace. Sometimes you can break and fix things without anyone even knowing which is very valuable if breaking something had some real risk to you.
Our backend is failing, it's on us-east-1 using AWS Lambda, Api Gateway, S3
Even this little tidbit is a bit of a wtf for me. Why do they consider it ok to have anything hosted in a single region?
At a different (unnamed) FAANG, we considered it unacceptable to have anything depend on a single region. Even the dinky little volunteer-run thing which ran https://internal.site.example/~someEngineer was expected to be multi-region, and was, because there was enough infrastructure for making things multi-region that it was usually pretty easy.
It was also greatly affecting Amazon.com itself. I kept getting sporadic 404 pages and one was during a purchase. Purchase history wasn't showing the product as purchased and I didn't receive an email, so I repurchased. Still no email, but the purchase didn't end in a 404, but the product still didn't show up in my purchase history. I have no idea if I purchased anything, or not. I have never had an issue purchasing. Normally get a confirmation email within 2 or so minutes and the sale is immediately reflected in purchase history. I was unaware of the greater problem at that moment or I would have steered clear at the first 404.
Part of me wonders how much they're actually going to pay out, given that their own status page has only indicated five services with moderate ("Increased API Error Rates") disruptions in service.
Either way—overt lies or engineering incompetence—it’s disappointing!
Wouldn't affected companies be incentivized to make a lawsuit about AMZ lying about status? It would be easy to prove and costly to defend from AWS standpoint.
Every AWS customer has a personal health dashboard that links the issues to their services which is updated much faster, and links issues to your affected resources. Additionally requests for credits are done by the customer service team who have even more information.
Two is one & one is none.
I think it only stopped when the storage services got to the "deprecated, and we're not bothering to do a failover because dependent teams who care should just use something else, because this one is being shut down any year now". (I don't agree with that decision, obviously ;) but I do have sympathy for the team stuck running a condemned service. Sigh.)
After stuff was migrated to the new storage service (probably somewhere in the 2017-2019 range but I have no idea when), I have no idea how DR/failover worked.
How long before Meta takes over for Facebook?
Especially when you are getting burned by an outage.
I'm guessing Google, on the basis of the recently published (to the public) "I just want to serve 5TB"[1] video. If it isn't Google, then the broccoli man video is still a cogent reminder that unyielding multi-region rigor comes with costs.
The point is that we made it easier. By the time I left, things were basically just multi-region by default. (To be sure, there were still sharp edges. Services which needed to store data (like, databases) were a nightmare to manage. Services which needed to be in the same region specific instances of other services, e.g. something which wanted to be running in the same region as wherever the master shard of its database was running, were another nasty case.)
The point was that every services was expected to be multi-region, which was enforced by regular fire drills, and if you didn't have a pretty darn good story about why regular announced downtime was fine, people would be asking serious questions.
And anything external going down for more than a minute or two (e.g. for a failover) would be inexcusable. Especially for something like a bloody login page.
They're cheap. HA is for their customers to pay more, not for Amazon which often lies during major outages. They would lose money on HA and they would lose money on acknowledging downtimes. They will lie as long as they benefit from it.
There are also various media articles but I can't tell which ones have significant new information beyond "outage".
Sagemaker is not working, I can't get to my work (notebook instance is frozen upon launch, with zero way to stop it or restart it) and Sagemaker Studio is also broken right now.
The length of this outage has blown my mind.
Rather, you use AWS because when it is down, it's down for everybody else as well. (Or at least they can nod their head in sympathy for the transient flakiness everybody experiences.) Then it comes back up and everybody forgets about the outage like it was just background noise. This is what's meant by "nobody ever got fired for buying (IBM|Microsoft)". The point is that when those products failed, you wouldn't get blamed for making that choice; in their time they were the one choice everybody excused even when it was an objectively poor choice.
As for me, I prefer hosting all my own stuff. My e-mail uptime is better than GMail, for example. However, when it is down or mail does bounce, I can't pass the buck.
Whenever his projects went down, he fought tooth and nail against any suggestion to update the status page. When forced to update the status page, he'd follow up with an extremely long "post-mortem" document that was really just a long winded explanation about why the outage was someone else's fault.
He later explained that in his department at Amazon, being at fault for an outage was one of the worst things that could happen to you. He wanted to avoid that mark any way possible.
YMMV, of course. Amazon is a big company and I've had other friends work there in different departments who said this wasn't common at all. I will always remember the look of sheer panic he had when we insisted that he update the status page to accurately reflect an outage, though.
And sure enough, PragmaticPulp did post a similar comment on a thread about Amazon India's alleged hire-to-fire policy 6 months back: https://news.ycombinator.com/item?id=27570411
You and I, we aren't among the 10000, but there are potentially 10000 others who might be: https://xkcd.com/1053/
It strikes a nerve with me because it caused so much trouble for everyone around him. He had other personal issues, though, so I should probably clarify that I'm not entirely blaming Amazon for his habits. Though his time at Amazon clearly did exacerbate his personal issues.
When you get to the level you want, you get to not really give a shit and actually do The Right Thing. However, for all of the engineers clamoring to get out of the intermediate brick laying trenches, opening an incident can have pervasive incentives.
We may run into the problem of everything documented and possible deliberate acts but for a service that relies heavily on uptime, that’s a small price to pay for a bulletproof operation.
Post-mortems can sometime be thought of like safety training. There is a big imbalance of time dedicated to learning proper safety handling just for those small incidences.
It make AWS engineers look stupid, because it looks like they are not monitoring their services.
Management.
No-blame analysis is a much better pattern. Everyone wins. It's about building the system that builds the system. Stuff broke; fix the stuff that broke, then fix the things that let stuff break.
I realize that isn't an easy thing to do. Often the best bet is to just jump around till you find a company that isn't a cultural superfund site.
Yea, except it doesn't work in practice. I work with a lot of people who come from places with "blameless" post-mortem 'culture' and they've evangelized such a thing extensively.
You know what all those people have proven themselves to really excel at? Blaming people.
It's like saying unit tests don't work in practice because bugs got through.
In other words, “no-blame” should be an emergent property of a culture of trust. It’s not something you can prescribe.
It's like a seed for crystal growth. Small company is exactly the best time to implement these things, because other employees will try to match the cultural norms and habits.
The fact that I've gotten it to the point of using git with automated build and deployment is a small miracle in itself. Not everybody gets to start from a clean slate.
Engineers should be experts and you should be able to trust them to make reasonable choices about the management of their projects.
That doesn't mean there can't be some checks in place, and it doesn't mean that all engineers should be perfect.
But you also have to acknowledge that adding all of those safeties has a cost. You can be a competent person who requires fewer safeties or less competent with more safeties.
Which one provides more value to an organization?
Neither, they both provide the same value in the long term.
Senior engineers cannot execute on everything they commit to without having a team of engineers they work with. If nobody trains junior engineers, the discipline would go extinct.
Senior engineers provide value by building guardrails to enable junior engineers to provide value by delivering with more confidence.
network_cli remove_routes [--region us-east-1]
Blaming the operator that they should have known that running network_cli remove_routes
will take down all regions because the region wasn't specified is the kind of thing as to what's being called out here.All of the tools need to not default to breaking the world. That is the first and foremost thing being pushed. If an engineer is remotely afraid to come forwards (beyond self-shame/judgement) after an incident, and say "hey, I accidentally did this thing", then the situation will never get any better.
That doesn't mean that engineers don't have the ability to break things, but it means it's harder (and very intentionally so) for a stressed out human operator to do the wrong thing by accident. Accidents happen. Do you just plan on never getting into a car accident, or do you wear a seat belt?
-You don't hide anything
-Errors will be made
-After training/mission everyone talks about the errors (or potential ones) and how to prevent them
-You don't make the same error twice
Being afraid to make errors and learn from them creates a culture of hiding, a culture of denial and especially being afraid to take responsibility.
But usually it isn't the same person making the same mistake, usually it is someone else making the same mistake and nobody thought of updating processes/documentation to the point that the error would have been caught in time. Maybe they'll fix that after the second time ;)
"Yes, John has made mistakes and he's always copped to them immediately and worked to prevent them from happening again in the future. You know who doesn't make mistakes? People who don't do anything."
People spend all this time threat modelling their stuff against malefactors, and yet so often people don't spend any time thinking about the threat model of decay. They don't do it adding new dependencies (build- or runtime), and therefore are unprepared to handle an outage.
There's a good reason for this, of course: modern software "best practices" encourage moving fast and breaking things, which includes "add this dependency we know nothing about, and which gives an unknown entity the power to poison our code or take down our service, arbitrarily, at runtime, but hey its a cool thing with lots of github stars and it's only one 'npm install' away".
Just want to end with this PSA: Dependencies bad.
Yes
> Did I lack due diligence in choosing to accept the risk that the other team couldn't deliver?
Yes
That doesn't extend to ridiculous lengths but as a rule you should engineer around any single point of failure.
If its standard CPU that everybody else uses, and its not known to be bad then no.
Same for software. Is it ok to have dependency on AWS services ? Their history shows yes. Dependency on brand new SaaS product ? Nothing mission critical.
Or npm/crates/pip packages. Packages that have been around and seedily maintained for few years, have active users, are worth checking out. Some random project from single developer ? Consider vendoring (and owning if necessary ) it.
Now say you're a medium sized hyper-growth company in a competitive space. Does spending 10 times more and waiting 5 times longer for redundancy make business sense? You could argue that it'd be irresponsible to over-engineer the system in this case, since you delay getting your product out and potentially lose $ and ground to competitors.
I don't think a black and white "yes, you should be punished" view is productive here.
Taking blame is a purely punitive action and solves nothing. Taking responsibility means it's your job to correct the problem.
I find that the more "political" the culture in the organization is, the more likely everyone is to search for a scapegoat to protect their own image when a mistake happens. The higher you go up in the management chain, the more important vanity becomes, and the more you see it happening.
I have made plenty of technical decisions that turned out to be the wrong call in retrospect. I took _responsibility_ for those by learning from the mistake and reversing or fixing whatever was implemented. However, I never willfully took _blame_ for those mistakes because I believed I was doing the best job I could at the time.
Likewise, the systems I manage sometimes fail because something that another team manages failed. Sometimes it's something dumb and could have easily been prevented. In these cases, it's easy point blame and say, "Not our fault! That team or that person is being a fuckup and causing our stuff to break!" It's harder but much more useful to reach out and say, "hey, I see x system isn't doing what we expect, can we work together to fix it?"
People tend to believe that if you can describe a problem that means you can prescribe a solution. Often times, the only way to survive is to make it clear that the first thing you are doing is describing the problem.
After you do that, and it's clear that's all you are doing, then you follow up with a prescriptive description where you place clearly what could be done to manage a future scenario.
If you don't create this bright line, you create a confused interpretation.
The injustice that can and does happen is that you're explicitly given a narrow responsibility during development, and then a much broader responsibility during operation. This is patently unfair, and very common. For something like a failed uService you want to blame "the architect" that didn't anticipate these system level failures. What is the solution? Have plan b (and plan c) ready to go. If these services don't exist, then you must build them. It also implies a level of indirection that most systems aren't comfortable with, because we want to consume services directly (and for good reason) but reliability requires that you never, ever consume a service directly, but instead from an in-process location that is failure aware.
This is why reliable software is hard, and engineers are expensive.
Oh, and it's also why you generally do NOT want to defer the last build step to runtime in the browser. If you start combining services on both the client and server, you're in for a world of hurt.
Remember: it may not be your fault, but it still is your problem.
You get hit by a car and injured. The accident is the other driver's fault, but getting to the ER is your problem. The other driver may help and call an ambulance, but they might not even be able to help you if they also got hurt in the car crash.
I love the ubiquity of thirdparty software from strangers, and the lack of bureaucratic gatekeepers. but I also hate it in ways. and not enough people know about the dangers of this second thing.
dat gap tho... which was my point. smart black hats will be exploiting this gap, at scale. and the strategy will work because the majority of folks seem to be either lazy, ignorant or simply hurried for time.
and btw your 1st sentence was rude. constructive feedback for the future
Ok, let's take an organization, let's call them, say Ammizzun. Totally not Amazon. Let's say you have a very aggressive hire/fire policy which worked really well in rapid scaling and growth of your company. Now you have a million odd customers highly dependent on systems that were built by people that are now one? two? three? four? hire/fire generations up-or-out or cashed-out cycles ago.
So.... who owns it if the people that wrote it are lllloooooonnnnggg gone? Like, not just long gone one or two cycles ago so some institutional memory exists. I mean, GONE.
They are now on the 6th butt in that seat in 4 years. That poor fellow is entirely blameless for the mess that accumulated over time.
Both times I was called out as a failure in my performance eval. Second time, I resigned and told them to find a better leader.
Happy now I am out of such shitty place.
Doesn't sound like it.
On another note, thanks for building some awesome stuff -- walmart.com is awesome. I have both Prime and whatever-they're-currently-calling Walmart's version and I love that Walmart doesn't appear to mix SKU's together in the same bin which seems to cause counterfeiting fraud at Amazon.
An amazon listing doesn't guarantee a particular SKU.
Although, Amazon is the worst, then Walmart (still much better than Amazon since you can at least filter). The others are not bad in my experience.
I think they got rid of it at some point though.
I've had push backs on my postmortems before because of phrasing that could be constituted as laying some of the blame on some person/team when it's supposed to be blameless.
And for a long time, it was fairly blameless. You would still be punished with the extra work of writing high quality postmortems, but I have seen people accidentally bring down critical tier-1 services and not be adversely affected in terms of promotion, etc.
But somewhere along the way, it became politicized. Things like the wheel of death, public grilling of teams on why they didn't follow one of the thousands of best practices, etc, etc. Some orgs are still pretty good at keeping it blameless at the individual level, but... being a big company, your mileage may vary.
The moment that affects promotions negatively, or your coworkers throw you under the bus, you should 1) be assertive and 2) proof-read your resume as a precursor to job hunting.
I wanted to be proactive and fix things before they became an issue, but such things just drained life out of me, to the point I just left.
If I think Tom has a toxic combination of poor judgement, Dunning-Kruger syndrome, and a hint of narcissism (I'm not sure but I may be repeating myself here), such that he won't listen to reason and he actively steers others into bad situations (and especially if he then disappears when shit hits the fan), then I will nail him to a fucking cross every chance I get. Public shaming is only a tool for getting people to discount advice from a bad actor. If it comes down to a vote between my idea and his, then I'm going to make sure everyone knows that his bets keep biting us in the ass. This guy kinda sounds like the Toxic Tom.
What is important when I turned out to be the cause of the issue is a bit like some court cases. Would a reasonable person in this situation have come to the same conclusion I did? If so, then I'm just the person who lost the lottery. Either way, fixing it for me might fix it for other people. Sometimes the answer is, "I was trying to juggle three things at once and a ball got dropped." If the process dictated those three things then the process is wrong, or the tooling is wrong. If someone was asking me questions we should think about being more pro-active about deflecting them to someone else or asking them to come back in a half hour. Or maybe I shouldn't be trying to watch training videos while babysitting a deployment to production.
If you never say "my bad" then your advice starts to sound like a lecture, and people avoid lectures so then you never get the whole story. Also as an engineer you should know that owning a mistake early on lets you get to what most of us consider the interesting bit of solving the problem instead of talking about feelings for an hour and then using whatever is left of your brain afterward to fix the problem. In fact in some cases you can shut down someone who is about to start a rant (which is funny as hell because they look like their head is about to pop like a balloon when you say, "yep, I broke it, let's move on to how do we fix it?")
"Blameless" to me means you acknowledge that the ultimate problem isn't that someone made a mistake that caused an outage. The problem is that you had a system in place where someone could make a single mistake and cause an outage.
If someone fat-fingers a SQL query and drops your database, the problem isn't that they need typing lessons! If you put a DBA in a position where they have to be typing SQL directly at a production DB to do their job, THAT is the cause of the outage, the actual DBA's error is almost irrelevant because it would have happened eventually to someone.
It's really easy for another developer to figure out who I'm talking about. Managers can't be arsed to figure it out, or at least pretend like they don't know.
It may also be that the cause is willful negligence, intentionally circumventing barriers for some personal reason.
And, of course, it may be that the cause is explicitly malicious (e.g. internal fraud, or the intent to sabotage someone) and at least part of the blame directly lies on the culprit, and not only on those who failed to notice and stop them.
Jeff himself has said many times in All Hands and in public "Amazon is the best place to fail". Mainly because things will break, it's not that they break that's interesting, it's what you've learned and how you can avoid that problem in the future.
With the size of your customer base there were man years spent confirming the outage after checking the status.
Imagine how stressful life would be thinking that you had to be perfect all the time.
Blame deflection is a recipe for repeat outages and unhappy customers.
Entirely possible, and something I've always suspected.
Some AWS managers and engineers bring their corporate cultural baggage with them when they join AWS and it takes a few years to unlearn it.
Amazon is a huge company so I have no doubt YMMV depending on your manager.
The truth (as always) is more complex:
* No, this isn't the broad culture. It's not even a blip. These are EXCEPTIONAL circumstances by extremely bad teams that - if and when found out - would be intervened dramatically.
* The broad culture is blameless post-mortems. Not whose fault is it. But what was the problem and how to fix it. And one of the internal "Ten commandments of AWS availability" is you own your dependencies. You don't blame others.
* Depending on the service one customer's experience is not the broad experience. Someone might be having a really bad day but 99.9% of the region is operating successfully, so there is no reason to update the overall status dashboard.
* Every AWS customer has a PERSONAL health dashboard in the console that should indicate their experience.
* Yes, VP approval is needed to make any updates on the status dashboard. But that's not as hard as it may seem. AWS executives are extremely operation-obsessed, and when there is an outage of any size are engaged with their service teams immediately.
The whole us-east-1 management console is gone, what is Amazon posting for the management console on their website?
"Service degradation"
It's not a degradation if it's outright down. Use the red status a little bit more often, this is a "disruption", not a "degradation".
An increase in error rates - no biggie, any large system is going to have errors. But when 80%+ of customers loads in the region are impacted (cross availability zones for whatever good those do) - that counts as down doesn't it? Error rates in one AZ - degraded. Multi-AZ failures - down?
So i aggree that degraded is not the proper wording - but it's / was not completly vanished. so.... hard to tell what is an common acceptable wording here.
To your point, for support center (which doesn't show a region) it says:
Description
Increased Error Rates
[09:01 AM PST] We are investigating increased error rates for the Support Center console and Support API in the US-EAST-1 Region.
[09:26 AM PST] We can confirm increased error rates for the Support Center console and Support API in the US-EAST-1 Region. We have identified the root cause of the issue and are working towards resolution.
It's only updated when a large percentage of customers are impacted, and most of the time this number is less than what the HN echo chamber makes it appear to be.
But if the official records say everything is green, a customer is going to have to push a lot harder to get the credits. There is a massive incentivization to “stay green”.
If the console works 100% in us-east-2 and not in us-east-1 why would they put the console completely down in us-east?
And a cleaner is called a "floor technician".
Nothing really out of the ordinary for a service to be called degraded while "hey, the cache might still be working right?" ... or "Well you know, it works every other day except today, so it's just degradation" :-)
My experience generally aligns with amzn-throw, but this right here is why. There's a manual step here and there's always drama surrounding it. The process to update the status page is fully automated on both sides of this step, if you removed VP approval, the page would update immediately. So if the page doesn't update, it is always a VP dragging their feet. Even worse is that lags in this step were never discussed in the postmortem reviews that I was a part of.
Let's not pretend businesses haven't been intentionally advertising in deceitful ways for decades if not hundreds of years. This just happens to be current strategy in tech of lying and deceiving customers to limit liability, responsibility, and recourse actions.
To be fair, it's not it's not just Amazon, they just happen to be the largest and targeted whipping boys on the block. Few businesses under any circumstances will admit to liability under any circumstances. Liability has to always be assessed externally.
Agreed, teams should invest resources in architecting their systems in a way that can withstand broken dependencies. How does AWS teams account for "core" dependencies (e.g. auth) that may not have alternatives?
The key difference is the perspective. If reliability is bad that’s an organizational problem and blaming or punishing one engineer won’t fix that.
[1] An example ladder from Patreon: https://levels.patreon.com/
The key difference between what and what?
This was also towards the end of the golden age of Google, when the percentage of top talent was a lot higher.
How should their performance be evaluated, if not by the rote number of mistakes that can be pinned onto the person, and their combined impact? (Was that the question?)
Now i guess its possible that the 1000s and 1000s of us who noticed and commented are some tiny fraction of the user base but if thats so you could at least publish a follow up like other vendors do that says something like 0.00001% of API requests failed effecting an estimated 0.001% of our users at the time.
See, AWS is basically turning into a long standing utility that needs to be reliable.
Hey, do most institutions like that completely turn over their staff every three years? Yeah, no.
Great for building it out and grabbing market share.
Maybe not for being the basis of a reliable substrate of the modern internet.
If there are dozens of bespoke systems that keep AWS afloat (disclosure: I have friends who worked there, and there are, and also Conway's law), but if the people who wrote them are three generations of HIRE/FIRE ago....
Not good.
Maybe THEY will go to a COMPETITOR and THINGS MOVE ON if it's THAT BAD. I wasn't sure what the pattern for all caps was, so just giving it a shot there. Apologies if it's incorrect.
Everybody is very slow to update their outage pages because of SLAs. It's in a company's financial interest to deny outages and when they are undeniable to make them appear as short as possible. Status pages updating slowly is definitely by design.
There's no large dev platform I've used that this wasn't true of their status pages.
If services are clearly down, why is this needed ? I can understand the oversights required for a company like Amazon but this sounds strange to me. If services are clearly down, I want that damn status update right away as a customer.
You mean the one that is down right now?
It's not acceptable.
It reflects exactly my experience there.
Blameless post-mortem, stick to the facts and how the situation could be avoided/reduced/shortened/handled better for next time.
In fact, one of the guidelines for writing COE (Correction Of Error, Amazon's jargon for Post Mortem) is that you never mention names but use functions and if necessary teams involved:
1. Personal names don't mean anything except to the people who were there on the incident at the time. Someone reading the CoE on the other side of the world or 6 months from now won't understand who did what and why. 2. It stands in the way of honest accountability.
I love that. Build your service to be robust. Never assume that dependencies are 100% reliable. Gracefully handle failures. Don't just go hard down, or worse, sure horribly in a way that you can't recover from automatically when you're dependencies come back. I've seen a single database outage cause cascading failures across a whole site even though most services had no direct connection to the database. And recovery had to be done in order of dependency, or else you're playing whack-a-mole for an hour)
> VP approval is needed to make updates on the status board.
Isn't that normal? Updating the status has a cost (reparations to customers if you breach SLA). You don't want some on-call engineer stressing over the status page while trying to recover stuff.
In my experience, it's rarely clear who was at fault for any sort of non-trivial outage. The issue tends to be at interfaces and involve multiple owners.
I don't know if you're exaggerating or not, but even if true why would anyone show that emotion about losing a job in the worst case?
You certainly had a lot of relevate-to-todays-top-hn-post stories throughout you career. And I'm less and less surprised to continuously find PragmaticPulp as one of the top commenters if not the top that resonates with a good chunk of HN.
I guess 100% is technically an increase.
Issues started ~1:24am ET and resolved around 7:31am ET.
Then really kicked in at a much larger scale at 10:32am ET.
We're now seeing failures with connections to RDS Postgres and other services.
Console is completely unavailable to me.
First engineer found a clever hack using bubble gum
>Then really kicked in at a much larger scale at 10:32am ET.
Bubble gum dried out, and the connector lost connection again. Now, connector also fouled by the gum making a full replacement required.
Edit: Just to add, very simple binary automations are even possible without a central controller. Like, I have Insteon motion sensors that trigger a lighting scene when they detect motion. These are super simplistic though.
* Visit the console directly from another region's URL (e.g., https://us-east-2.console.aws.amazon.com/console/home?region...). You can try this after you've successfully signed in but see the console failing to load as well.
* If your AWS SSO app is hosted in a region other than us-east-1, you're probably fine to continue signing in with other accounts/roles.
Of course, if all your stuff is in us-east-1, you're out of luck.
EDIT: Removed incorrect advice about running AWS SSO in multiple regions.
Is this possible?
> AWS Organizations only supports one AWS SSO Region at a time. If you want to make AWS SSO available in a different Region, you must first delete your current AWS SSO configuration. Switching to a different Region also changes the URL for the user portal. [0]
This seems to indicate you can only have one region.
[0] https://docs.aws.amazon.com/singlesignon/latest/userguide/re...
We are incredibly blind to just trust just 3 cloud providers with the operational success of basically everything we do.
Why hasn't the industry come up with an alternative?
Most financial institutions are implementing their own clouds, I can't think of any major one that is reliant on public cloud to the extent transactions would stop.
>Why hasn't the industry come up with an alternative?
You mean like building datacenters and hosting your own gear?
The agreement is more of a hybrid cloud arrangement with AWS Outposts.
FTA:
>Core to Nasdaq’s move to AWS will be AWS Outposts, which extend AWS infrastructure, services, APIs, and tools to virtually any datacenter, co-location space, or on-premises facility. Nasdaq plans to incorporate AWS Outposts directly into its core network to deliver ultra-low-latency edge compute capabilities from its primary data center in Carteret, NJ.
They are also starting small, with Nasdaq MRX
This is much less about moving NASDAQ (or other exchanges) to be fully owned/maintained by Amazon, and more about wanting to take advantage of development tooling and resources and services AWS provides, but within the confines of an owned/maintained data center. I'm sure as this partnership grows, racks and racks will be in Amazon's data centers too, but this is a hybrid approach.
I would also bet a significant amount of money that when NASDAQ does go full "cloud" (or hybrid, as it were), it won't be in the same US-east region co-mingling with the rest of the consumer web, but with its own redundant services and connections and networking stack.
NASDAQ wants to modernize its infrastructure but it absolutely doesn't want to offload it to a cloud provider. That's why it's a hybrid partnership.
Companies do not change their whole strategy from a capex-driven traditional self-hosting environment to opex-driven cloud hosting because their IT people are lazy; it is typically an exec-level decision.
Medical devices, banks, the military, etc. should generally run on their own hardware. The next photo-sharing app? It's just not worth it until they hit tremendous scale.
On the second though, at some point, infrastructure like AWS are going to be more reliable than what many banks, medical device operators etc can provide themselves. asking them to stay on their own hardware is asking for that industry to remain slow, bespoke and expensive.
It is incredibly difficult for non-tech companies to hire quality software and infrastructure engineers - they usually pay less and the problems aren't as interesting.
vanishing or delayed six hours? I mean
Multiple-region redundancy costs more both in initial planning/setup as well as monthly fees so a lot of AWS customers choose to just not do it.
Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.
Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.
AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.
And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.
edit: I will add that AWS specifically is meh and I wouldn't use it myself, there's better IASS. But it's insanity to even imagine self-hosted is more reliable than using even the shittiest of IASS providers.
That might be true, but the effects of any given outage would be felt much less widely. If Disney has an outage, I can just find a movie on Netflix to watch instead. But now if one provider goes down, it can take down everything. To me, the problem isn't the cloud per se, it's one player's dominance in the space. We've taken the inherently distributed structure of the internet and re-centralized it, losing some robustness along the way.
If my system has an hour of downtime every year and the dozen other systems it interacts with and depends on each have an hour of downtime every year, it can be better that those tend to be correlated rather than independent.
If you're a company relying upon AWS for your business, is it okay if you're down for a day, or two while you wait for AWS to resolve it's issue?
Apple designed their own silicon, a third party manufactures and packages it for them.
You mean like that rule, pedant? It's not name calling if it's an accurate representation of one's behavior.
ped·ant /ˈpednt/: noun a person who is excessively concerned with minor details and rules or with displaying academic learning.
Apple designs the M1. But TSMC (and possibly Samsung) actually manufacture the chips.
Was a nightmare recovering data. Even when the service was operational was sub par.
Just saying perhaps the “shittiest” providers may not be more reliable.
It’s had two in 13 months
> Heck, why stop at having servers on-site? Cast your own silicon waffers, after all you don't want spectrum exploits.
That's an overblown argument. Nobody is saying that, but it's clear that businesses that maintain their own infrastructure would've avoided today's AWS' outage. So just avoiding a single level of abstraction would've kept your company running today.
> Because you are worst at it. If a specialist is this bad, and the market is fully open, then it's because the problem is hard.
The problem is hard mostly because of scale. If you're a small business running a few websites with a few million hits per month, it might be cheaper and easier to colocate a few servers and hire a few DevOps or old-school sysadmins to administer the infrastructure. The tooling is there, and is not much more difficult to manage than a hundred different AWS products. I'm actually more worried about the DevOps trend where engineers are trained purely on cloud infrastructure and don't understand low-level tooling these systems are built on.
> AWS has fewer outages in one zone alone than the best self-hosted institutions, your facebooks and petagons. In-house servers would lead to an insane amount of outage.
That's anecdotal and would depend on the capability of your DevOps team and your in-house / colocation facility.
> And guess what? AWS (and all other IAAS providers) will beg you to use multiple region because of this. The team/person that has millions of dollars a day staked on a single AWS region is an idiot and could not be entrusted to order a gaming PC from newegg, let alone run an in-house datacenter.
Oh great, so the solution is to put even more of our eggs in a single provider's basket? The real solution would be having failover to a different cloud provider, and the infrastructure changes needed for that are _far_ from trivial. Even with that, there's only 3 major cloud providers you can pick from. Again, colocation in a trusted datacenter would've avoided all of this.
What are good examples of
>a small business running a few websites with a few million hits per month, it might be cheaper and easier to colocate a few servers and hire a few DevOps or old-school sysadmins to administer the infrastructure.
and how often do they go down?
Heck, I use old T430 for my home server and still it doesn't go down on completely random occasions (but thats very simplified example, I know)
No idea what are the standards for other companies.
When Netflix was running its own datacenters in 2008, they had a 3 day outage from a database corruption and couldn't ship DVDs to customers. That was the disaster that pushed CEO Reed Hastings to get out of managing his own datacenters and migrate to AWS.
The flaw in the reasoning that running your own hardware would avoid today's outage is that it doesn't also consider the extra unplanned outages on other days because your homegrown IT team (especially at non-tech companies) isn't as skilled as the engineers working at AWS/GCP/Azure.
Sure, that's trivially obvious. But how many other outages would they have had instead because they aren't as experienced at running this sort of infrastructure as AWS is?
You seem to be arguing from the a priori assumption that rolling your own is inherently more stable than renting infra from AWS, without actually providing any justification for that assumption.
You also seem to be under the assumption that any amount of downtime is always unnacceptable, and worth spending large amounts of time and effort to avoid. For a lot of businesses systems going down for a few hours every once in a while just isn't a big deal, and is much more preferable than spending thousands more on cloud bills, or hiring more full time staff to ensure X 9s of uptime.
To be fair, I'm not saying never use cloud providers. If your systems require the complexity cloud providers simplify, and you operate at a scale where it would be prohibitively expensive to maintain yourself, by all means go with a cloud provider. But it's clear that not many companies are prepared for this type of failure, and protecting against it is not trivial to accomplish. Not to mention the conceptual overhead and knowledge required with dealing with the provider's specific products, APIs, etc. Whereas maintaining these systems yourself is transferrable across any datacenter.
Ovh, scaleway, online.net, azure, gcp, aws
That's one's I've used in production, I've heard of a dozen more including big names like HP and IBM, I assume they can match aws for the most part.
...
That being said I agree multi tenant is the way to go for reliability. But I was pointing out that in this case even the simple solution of multi region on one provider was not implemented by those affected.
...
As for running your own data center as a small company. I have done it, buying components building servers and all.
Expenses and ISP issues aside, I can't imagine using in house without at least a few outages a year for anywhere near the price of hiring a DevOps person to build a MT solution for you.
If you think you can you've either never tried doing it OR you are being severely underpaid for your job.
Competent teams to build and run reliable in house infrastructure exist, and they can get you SLA similar to multi region AWS or GC (aka 100% over the last 5 years)... But the price tag has 7 to 8 figures in it.
will they? because AWS still puts new stuff in us-east-1 before anywhere else, and there is often a LONG delay before those things go to other regions. there are many other examples of why people use us-east-1 so often, but it all boils down to this: AWS encourage everyone to use us-east-1 and discourage the use of other regions for the same reasons.
if they want to change how and where people deploy, they should change how they encourage it's customers to deploy.
my employer uses multi-region deployments where possible, and we can't do that anywhere nearly as much as we'd like because of limitations that AWS has chosen to have.
so if cloud providers want to encourage multi-region adoption, they need to stop discouraging and outright preventing it, first.
I also think that they've been making a lot of the default region dropdowns and such point to CMH (us-east-2) to get folks to migrate away from IAD. Your contention that they're encouraging people to use that region just don't ring true to me.
Come to think of it (far down the second page of comments): Why east?
Amazon is still mainly in Seattle, right? And Silicon Valley is in California. So one would have thought the high-tech hub both of Amazon and of the USA in general is still in the west, not east. So why us-east-1 before anywhere else, and not us-west-1?
I'm not sure how many aws services are easy to spawn at multiple regions
Doesn't help you if it what goes down is AWS global services on which you directly, or other AWS services, depend (which tend to be tied to US-east-1).
It's not AWS fault here, it's the companies', which assume that it will never be down. In-house servers also have outages, it's a very naive assumption to think that it'd be all better if all of those services were using their own servers.
Facebook doesn't use AWS and they were down for several hours a couple weeks ago, and that's because they have way better engineers than the average company, working on their infrastructure, exclusively.
However if youbare on AWS many of your competitors are down while you are down, so they can't takeover your business.
The cloud is the solution to self managed data centers. Their value proposition is appealing: Focus on your core business and let us handle infrastructure for you.
This fits the needs of most small and medium sized businesses, there's no reason not to use the cloud and spend time and money on building and operating private data centers when the (perceived) chances of outages are so small.
Then, companies grow to a certain size where the benefits of having a self managed data center begins to outweight not having one. But at this point this becomes more of a strategic/political decision than merely a technical one, so it's not an easy shift.
There is a middle ground of bare-metal hosting. Abstract away the hardware and networking. Do the rest yourself.
If you would like to go back to the days of managing your own machines, be my guest. Remember those machines also live somewhere and were/are subject to the same BGP and routing issues we've seen over the past couple of years.
Personally, I'll deal with outages a few times a year for the peace of mind that there's a group of really talented people looking into for me.
If the companies at the scale you are talking about do not have multi-region and multi service (aws to azure for example) failover that's their fault, and nobody else's.
We used to have that, some companies still have the capability and know-how to build and run infrastructure that is reliable, distributed across many hosting providers before "cloud" became the "norm", but it goes along with "use or lose it".
"We have identified the root cause of the issues in the US-EAST-1 Region, which is a network issue with some network devices in that Region which is affecting multiple services, including the console but also services like S3. We are actively working towards recovery."
"All teams are engaged and continuing to work towards mitigation. We have confirmed the issue is due to multiple impaired network devices in the US-EAST-1 Region."
Doesn't sound like they are having a good day there!
We’ve been told to disable all pipelines even if we have time blockers or manual approval steps or failing tests
Some 3600 people have hit that button in the last ~15 minutes.
When Facebook's properties all went down in October, people were saying that AT&T and other cell phone carriers were also down - because they couldn't connect to FB/Insta/etc. There were even some media reports that cited Downdetector, seeming without understanding that they are basically crowdsourced and sometimes the crowd is wrong.
As of right now, this thread and one update from a twitter user, https://twitter.com/SiteRelEnby/status/1468253604876333059 are all we have. I went into disaster recovery mode when I saw our traffic dropped to 0 suddenly at 10:30am ET. That was just the SQS/something else preventing our ELB logs from being extracted to DataDog though.
It's a joke. Each time AWS/Azure/GCP is down their status page says all is fine.
For example, the sign-up page appears to be working: https://portal.aws.amazon.com/billing/signup#/start
Are websites that run on AWS us-east up? Are the AWS CLIs working?
Looks like actual services.
I cannot access the admin console.
$ aws cloudfront list-distributions
An error occurred (HttpTimeoutException) when calling the ListDistributions operation: Could not resolve DNS within remaining TTL of 4999 ms
However I can still access the distribution fineHowever, my Alexa (Echo) won't control my thermostat right now.
And my Ring app won't bring up my cameras.
Those services are run on AWS.
Virginia is for lovers, Ohio is for availability.
Multi-provider dns is a solved problem.
I got two 9 problems cuz of us-east-1
I left my two nines problems in us-east-1
If your company is shoehorning you into using multiple clouds and learning a dozen products, IAM and CICD dialects simultaneously because "being cloud dependent is bad", I feel bad for you.
Doing one cloud correctly from a current DevSecOps perspective is a multi-year ask. I estimate it takes about 25 people working full time on managing and securing infrastructure per cloud, minimum. This does not include certain matrixed people from legacy network/IAM teams. If you have the people, go for it.
Example: Payment/Administrative issues, rogue employee with access, deprecated service, inter-region routing issues, root certificate compromises... the list goes on and it is certainly not limited to single AZ.
A very good example, is that regardless of which of the 85 AZs you are in at aws, you are affected by this issue right now.
Multi-cloud with the right tooling is trivial. Investing in learning cloud-proprietary stacks is a waste of your investment. You're a clown if you think you need 25 people internally per cloud is required to "do it right".
There is no such thing as trivially setting up a secure, fully automated cloud stack, much less anything like a streamlined cloud agnostic toolset.
Deprecated services are not the discussion here. We're talking tactical availability, not strategic tools etc.
Rogue employees with access? You mean at the cloud provider or at your company? Still doesn't make sense. Cloud IAM is very difficult in large organizations, and each cloud does things differently.
I worked at fortune 100 finance on cloud security. Some things were quite dysfunctional, but the struggles and technical challenges are real and complex at a large organization. Perhaps you're working on a 50 employee greenfield startup. I'll hesitate to call you a clown as you did me, because that would be rude and dismissive of your experience (if any) in the field.
K8s is one of a hundred technologies to learn and use, and just because each cloud is supported by terraform, you can't swap a GCP terraform writer over to Azure in a day.
And no bank is without their uncloudable components.
With God, all things are possibleWhile my heroku apps are currently up, I am unable to push new versions.
Logging in to heroku dashboard (which does work), there is a message pointing to this heroku status incident for "Availability issues with upstream provider in the US region": https://status.heroku.com/incidents/2390
How can there be an outage severe enough to be effecting middleman customers like heroku, but the AWS status page is still all green?!?!
If whoever runs the AWS status page isn't embaressed, they really ought to be.
I advise not touching your Heroku setup right now. Even something like trying to restart a dyno might mean it doesn't come back since the slug is probably stored on S3 and that will fail.
I'm guessing other IoT things suffer from this same short sitedness as well.
Whats a graceful fallback? Switching to another hosting service when AWS goes down? Wouldn't that present another set of complications for a very small edge case at huge cost?
Rather than implement a dynamically switching backup in the event of AWS going down which is not trivial.
One has to crunch the numbers. What does a service outage cost your business every minute/hour/day/etc in terms of lost revenue, reputational damage, violated SLAs, and other factors? For some enterprises, it's well worth the added expense and trouble of having multi-site active-active setups that span clouds and on-prem.
If your product requires 100% uptime, you may need to look at backup options or design your product in such a way that can handle temporary cloud failures.
They write the videos to GCS storage in Google Cloud, and to S3 in AWS. Every point of their workflows are checkpointed and cross referenced across GCP and AWS. If either side drops the ball, the other picks it up.
So yes, you can design a super fault tolerant system. This company did it because failing to deliver a few ads would mean lose of major contracts.
Maybe I'm too old, but I can't imagine a seasoned dev, much less a tech lead, omitting planning for that failure mode
A constitutional property of a network is it's volatility. Nodes may fail. Edges may. You may not. Or you may. But then you're delivering no reliabilty but crap. Nice sunshine crap, maybe.
It's a little harder than blocking the DNS unfortunately. But nonetheless it always brings a smile to my face to see that there's a FOSS frontier for everything.
Building fallbacks require work. How much extra effort and overhead is needed to build something like this ? Sometimes the cost vs benefits says that it is ok not to do it. If AWS has an outage like this once a year, maybe we can deal with it (unless you are working with mission critical apps).
There is a societal resilience benefit to not having unnecessary cloud dependencies beyond the privacy stuff. It makes your society and economy more robust if you can continue in the face of remote failures/errors.
It is December 7th, after all.
Haha, it would be funny if the IC reaches out to BigTech when failures occur to let them know they need not be worried about data loses. They can just borrow a copy of the data IC is siphoning off them. /s?
This comment is how I know you don't work in the public sector. Those agencies' infrastructures are essentially run by contractors with a few GS personnel making bad decisions every chance they get and a few DoD personnel acting like their rank can fix technical problems.
So at least Amazon retail is feeling some of the pain from this outage!
If that's the case, it's 100% a feature. They want as little public proof of an outage after it's over and to put the burden on customers completely to prove they violated SLAs.
Scratch that... console is not loading at all now :)
Step 1: deploy status checks to an external cloud.
That being said, AWS status pages are up.
I don't think a region being down is something that you can be unsure about.
This is what Amazon, the startup, understood.
Step 1: Always make it right and make the customer happy, even if it hurts in $.
Step 2: If you find you're losing too much money over a particular issue, fix the issue.
Amazon, one of the world's largest companies, seems to have forgotten that the risk of not reporting accurately isn't money, but breaking the feedback chain. Once you start gaming metrics, no leaders know what's really important to work on internally, because no leaders know what the actual issues are. It's late Soviet Union in a nutshell. If everyone is gaming the system at all levels, then eventually the ability to objectively execute decreases, because effort is misallocated due to misunderstanding.
How come an action of a private company in a capitalist country is like the Soviet Union?
Hopefully it results in a class action lawsuit for enough money that Amazon decides that an automated system is better than trying to supply human judgement.
Random aside: any chance you are related to the Calculus on Manifolds Spivak?
He says that any good theorem is worth generalizing, and I've generalized that to any life rule.
looks like only four 9's
> looks like only four 9's
That's why the Germans are such good engineers. Did the drives fail? Nein.
Did the CPU overheat? Nein.
Did the power get cut? Nein.
Did the network go down? Nein.
That's "four neins" right there. 5N = 99.999%
3N = 99.9%
1N5 = 95%
5N is <43m12s downtime per month.It is very unlikely that Amazon would deliberately make your messages cross the Atlantic just to find an American region that is unable to serve you.
"We are experiencing an error. Our apologies – We will be back up soon."
Part of the reason is because their SLAs are based on that dashboard, and that dashboard going red has a financial cost to AWS, so like any financial cost, it needs approval.
Partial availability isn’t the same as no availability.
let me repeat that: my AWS trainign that is run by AWS that I pay AWS for isn't working, because AWS is having control plane (or other) issues. This is several hours after the initial incident. We're doing training in us-west-2, but the identity service and other components run in us-east-1.
edit: partly down; it's sporadically failing
i'd say when other companies who run their infrastruture on AWS are going out, it's hard to argue it's not a real outage.
But AWS status _has_ changed to yellow at this point. Probably heroku could be completely down because of an AWS problem, and AWS status would still not show red. But at least yellow tells us there's a problem, the distinction between yellow and red probably only matters at this point to lawyers arguing about the AWS SLA, the rest of us know yellow means "problems", red will never be seen, and green means "maybe problems anyway".
I believe the entire us-east-1 could be entirely missing, and they'd still only put a yellow not a red on status page. After all, the other regions are all fine, right?
Versus, if a company has X SLA contracts signed, that point to Y reimbursement for being out for Z minutes, so it's easily calculable.
Taken literally what you are saying is the service could be down and an executive could override that, preventing them for paying customers for a service outage, even if the service did have an outage and the customer could prove it (screenshots, metrics from other cloud providers, many different folks see it).
I'm sure there is some subtlety to this, but it does mean that large corps with influence should be talking to AWS to ensure that status information corresponds with actual service outages.
2. Cost is not an issue (until it is but you’re already locked in so oh well)
3. Faang has drained the talent pool of people who know how
If this needed a CEO to eventually get around to pressing a button that said "show users the actual information about a problem" that reflects poorly on amazon.
I think you mean "start thinking they can get what they pay for"
It’s very frustrating. Why even have them?
We can't even update our product to say it's down, because accessing the product requires a process that is currently dead.
> AWS Management Console Home page is currently unavailable.
> You can monitor status on the AWS Service Health Dashboard.
"AWS Service Health Dashboard" is a link to status.aws.amazon.com... which is ALL GREEN. So... thanks for the suggestion?
At this point the AWS service health dashboard is kind of famous for always been green isn't it? It's a joke to it's users. Do the folks who work on the relevant AWS internal team(s) know this, and just not have the resources to do anything about it, or what? If it's a harder problem than you'd think for interesting technical reasons, that'd be interesting to hear about.
Amazing and scary to see all the unrelated services down right now.
My favorite status page, though, is Slack's. You can read an article in the New York Times about how Slack was down for most of a day, and the status page is just like "some percentage of users experienced minor connectivity issues". "Some percentage" is code for "100%" and "minor" is code for "total". Good try.
Mom: Alexa, did you break something?
Alexa: No.
M: Really? What's this? 500 Internal server error
A: ok maybe management console is down
M: Anything else?
A: ...
A: ... ok maybe cloudwatch logs
M: Ah hah. What else?
A: That's it, I swear!
M: 503 ClientError
A: ...well okay secretsmanager might be busted too...
Me: Alexa, is AWS down right now?
Alexa: I'd rather not answer that
That's a bit like involving your kid in an argument between parents.
So what's a good health check actually report these days? Is it just about its own status, or should it include a breakdown of the status of external dependencies as part of its folded up status?
When the biggest cloud provider in the world is famous for gaslighting, it sets expectations for our whole industry.
It's fucking disgraceful that they tolerate such a lack of integrity in their organization.
$ 16:04 up 46 days, 7:02, 9 users, load averages: 3.68 3.56 3.18
US East 1 was down just over a year ago
https://www.theregister.com/2020/11/25/aws_down/
Meanwhile I moved one of my two internal DNS servers to a second site on 11 Nov 2020, and it's been up since then. One of my monitoring machines has been filling, rotating and deleting logs for 1,712 days with a load average in the c. 40 range for that whole time, just works.
If only there was a way to run stuff with an uptime of 364 days a year without using the cloud /s
(Also, OpEx vs CapEx financial shenanigans...)
All the same, I don't disagree with your point.
When it's down, it's my problem, and I can't do anything about it other than explain why I have no idea the system is broken and can't do anything about it.
"Why is my dohicky down? When will it be back?"
"Because it's raining, no idea"
May be accurate, it's also of no use.
But yes, Opex vs Capex, of course that's why you can lease your servers. It's far easier to spend company money with another $500 a month on AWS than spend $500 a year for a new machine.
In the cloud you're at the mercy of someone who doesn't even know you exist to fix it, without the protections that say an electric company has with supplying domestic users.
This thread has people unable to turn their lights on[0], it's hilarious how people tie their stuff to dependencies that aren't needed, with a history of constant failure.
If you want to host millions of people, then presumably your infrastructure can cope with the loss of a single AZ (and ideally the loss of Amazon as a whole). The vast majority of people will be far better off without their critical infrastructure going down in the middle of the day in the busiest sales season going.
Most people don't need to scale to a billion users overnight.
And what's a workday anyway, surely you operate globally?
The main application I run operates in 7 countries globally, but the US is the only one that has enough usage to require additional capacity during the workday. So out of 720 hours in a 30 day month, cloud scaling allows me to pay for additional capacity for only the (roughly) 160 hours that it's actually needed. It's a significant cost saver.
And because the scaling is based on actual metrics, it won't scale up on a holiday when nobody is using the application. More cost savings.
Obviously what I'm saying will not apply to all use cases, but I'm only talking about mine.
https://www.thomas-krenn.com/en/wiki/Processor_P-states_and_...
Which is an implementation of:
We first noticed failures because a tester happened to be testing in an env that uses the Amazon Pay sandbox.
I checked the prod site, and it wouldn't even ask me to login.
When I tried to login to SellerCentral to file a ticket - it told me my password (from a pw manager) was wrong. When I tried to reset, the OTP was ridiculously slow. Clicking "resend OTP" gives a "the OTP is incorrect" error message. When I finally got an OTP and put it in, the resulting page was a generic Amazon "404 page not found".
A while later, my original SellerCentral password, still un-changed because I never got another OTP to reset it, worked.
What the fuck kind of failure mode is that "services are down, so password must be wrong".
"[4:35 PM PST] With the network device issues resolved, we are not working towards recovery of any impaired services. We will provide additional updates for impaired services within the appropriate entry in the Service Health Dashboard."
I guess they gave up.
You don't want your service to go down, plus your team's comms at the same time.
https://slack.com/blog/news/slack-aws-drive-development-agil...
[1] https://www.theverge.com/2020/6/4/21280829/slack-amazon-aws-...
Edit: to be clear this is because I’m utterly helplessly unable to do anything at the moment.
To be fair when it’s AWS when something goes snap it’s not my problem which I’m happy about (until some wise ass at AWS hires me) :)
That's what I'm saying: you host it yourself in facilities owned by your company if you're not willing to have everyone twiddle their thumbs during this sort of event. Your DR environment can be co-located or hosted elsewhere.
Amazon customer service can't help me handle an order.
It us-east-1 is out, then half of Amazon is out too.
They probably paint themselves in a corner just like facebook few weeks ago.
This make me think;
Could it be that one day the internet will have a total global outage and it will take few days to recover?
That might be tricky if they are remote and not on the same AS as their router access points and have no completely out of band access, but you're still talking hours at most.
And then Health Dashboard is 100% green. What a joke.
Status pages are virtually useless these days
Every major company has moved away from having accurate status pages.
So unless regulation gets implemented that says otherwise, there's zero incentive for any company to maintain an accurate status page.
If something I'm responsible goes down to the point that my stakeholders are complaining (which is something seriously wrong), they are not going to be happy with "oh the cloud was down, not my fault"
If AWS is down or not it meaningless to me, if my service running on AWS is down or not is the key metric.
If a service is down and I can't get into it, then chatter on things like outages mailing list, or HN, will let me know if it's yet another cloud failure, or if it's something that's affecting my machine only.
Right, there should be an "alliance" of customers from different large providers (something like a Union but instead of workers, it would be customers). They are the ones that should measure SLAs and hold the provider accountable.
If not happy with the results switch.
But I'm leery of any business who's so dishonest they fear any outside oversight that brings repercussions for said dishonesty.
"If not happy, switch" is silly - it's not the customer's problem. And if you're a large customer and have invested heavily in getting staff trained on AWS, you can't just move.
Really not that complicated.
If I find out this is why I couldn't get my Happy Meal this morning I'm going to be really, really grumpy.
EDIT: I'm REALLY grumpy now:
https://aws.amazon.com/blogs/industries/aws-is-how-mcdonalds...
What am I supposed to do for lunch now? Go to the drive through and order like a normal person? /s
Grumble grumble
I know it's cheap but seriously... not worth it. Many of us have the scars to prove this.
https://docs.aws.amazon.com/general/latest/gr/ao.html
As a result, even if you weren’t using that region but you were using the API you were hosed for 6+ hours, And the status page never acknowledged that it was out of action.
We had the following Terraform in our production pipeline:
data “aws_organizations_organization” “current” {}
As a result, all of our deployments to our EU regions were borked. Of course, we couldn’t raise a support case because the support system was also down, and despite escalating to our TAM weren’t able to get the status page to reflect reality.
My concern is that the Organizations API specifically will be brushed under the carpet and we will still have a single point of failure in a region which we never intend to use.
Can't login to AWS console at signin.aws.amazon.com:
Unable to execute HTTP request: sts.us-east-1.amazonaws.com. Please try again.----
The answer likely depends on the specific thing, but I'd argue most people would take the better version of something at the risk of it not working 1 or 2 days per year.
According to an internal message I saw, their monitoring stuff is fucked too.
It does tack me to a non-region specific login page which is up, and then redirects back to us-west-2 which works.
If I go to my bookmarked login page it breaks because, presumably, it hits something that is broken in us-east-1.
D'oh!
Error 503
We're sorry, something went wrong.
Please try again...wait...wait...yep, try reload/refresh now.
But if you are seeing this again, please report it here.
Please explain which page you were at and where on it that you clicked
Thank you!
This also confirms it: https://downdetector.com/status/imdb/
Connection was closed before we received a valid response from endpoint URL: "https://codepipeline.us-east-1.amazonaws.com/".
We’ve been told to manually disable them to ensure integrity of our services when it recovers
https://twitter.com/amontalenti/status/1468265799458639877
Segment is publicly reporting issues delivering to Firehose, and one of my company's real-time monitors also triggered for Kinesis Firehose an hour ago.
Update:
By my sniff of it, some “core” APIs are down for S3 and EC2 (e.g. GET/PUT on S3 and node create/delete on EC2). Systems like Kinesis Firehose and DynamoDB rely on these APIs under the hood (“serverless” is just “a server in someone else’s data center”).
Further update:
There is a workaround available for the AWS Console login issue. You can use https://us-west-2.console.aws.amazon.com/ to get in -- it's just the landing page that is down (because the landing page is in the affected region).
I'm not complaining, I enjoyed the nostaliga - sometimes the web still feels like the late 90s
> Firehose error: Slow down.
I wonder if they had to introduce a new exception / rate-limit to mediate this issue ...
At least the updates are amusing:
"9:18 AM PST We can confirm degraded Contact handling by agents in the US-EAST-1 Region. Agents may experience issues logging in or being connected with end-customers."
WTF is "contact handling", an "agent" or an "end-customer"?
How about something like "We are confirming that some users are not able to connect to AWS services in us-east-1. We're looking into it."
Nothing on their status page. But the Console is not working.
- writing to Firehose (S3-backed)
- publishing to eventbridge
- terraform commands to ECS' API are stuck/hanging
Other spurious errors involving kinesis but nothing alarming. us-east-1
I am quite sure the AWS support will be getting many refund requests over the course of the week.
They need a monitor for the monitoring.
Solution: push it to production on the zone with the most users and see what breaks.
If the team follows pipeline best practices, they are supposed to deploy to a single small region first, wait 24 hours, and then deploy to more, wait more, and deploy to more, until finally deploying to us-east-1.
Status page shows all green but AWS confirms differently when on a customer phone call.
The us-east-1 region is consistently pushing the limits of scale for the AWS services, thus is has way more problems than other regions.
494 ERROR and "We're sorry Something went wrong with our website, please try again later."
Edit: wow that webpage is humongous... never heard of paging?
Wonder if crime rates might eventually spike up if aws goes down, in an utopian world where Amazon gets everyone to use ring.
"AWS is down! Christmas came early boys! Roll out..."
Seems to be the trend with the last 5-6 big cloud outages.
Internal Error
Please try again laterThat's the error I am getting when logging to aws console
WTF
Looks like it might be getting worse
error logging into aws role using saml assertion: error retrieving STS credentials using SAML: ServiceUnavailable: status code: 503
Basically Amazon fucked all their own products too.
Also isolation is not as good as they would have you believe: I am unable to login to AWS Quicksight in us-west-2...
Edit: Is this the longest running broad outage for AWS yet?
Additionally, for complex apps, automatic cross-region disaster recovery can take tens or even hundreds of dev years, something most small to midsized companies can't afford.
This one is pretty epic (pun intended). Bad enough that Down Detector [0] shows "Reports indicate there may be a widespread outage at Amazon Web Services, which may be impacting your service." in a red alert bar at the top.
That's the part I find interesting.