Goddamn AWS seems to be down again 4th time this month alone
downdetector.in
downdetector.in
I feel obligated to point out that this is a very editorialized title, which is against site guidelines. Not... that I don't sympathize, just that it's perhaps a bit out of line even for providing context.
Over 200 services, 86 availability zones and 26 regions. We might as well post every 5 minutes if we post every time something on AWS goes down. And yes AWS is more dependent on us-east-1 but one of the outrages posts was for a minor outage in us-west-1.
By analogy cars crash every few minutes in my city (Las Vegas), but I can’t conclude that indicates some inherent (or easily fixable) problem with the road infrastructure or the design of cars, or the skill of drivers. The frequent crashes rarely affect me.
It was both us-west-1 and us-west-2, for 45 minutes or so. Was interesting for Auth0 customers, as they use us-west-2 as their primary, and us-west-1 is the failover location.
and I'm averaging six nines uptime over the past year on a $5 KVM virtual machine hosted at a sketchy hosting company, but that doesn't mean something catastrophic can't or won't happen unexpectedly
note that I count the six nines as unplanned downtime, it's had less availability than that because I reboot it for a newer kernel and debian-stable system updates maybe every 4-5 months.
Imagine Comcast or Verizon fucking up their configs, you might not be able to access AWS stuff but that doesn’t mean that AWS is down.
But hey, the real story wouldn’t make the front page, so lets just stick to the fabricated narrative.
1. Playing the game yourself.
2. Watching other people play.
3. Talking about people playing the game.
4. Listening to someone else talk about people playing the game (i.e. sports commentators).
5. Talking about what sports commentators have to say.
In sports and in threads like this we're mostly at 3, 4, or 5. Even people who work at AWS posting here may only be "watching," i.e. working with second-hand information. While it's entertaining to nerds to speculate about root causes and possible solutions, and the dire predicted fallout from problems (especially when they happen to companies they dislike), it's just chatter and doesn't actually address any real issue, or inform any decisions that mean anything.
Say that you're answering the on-call pager, but with "elevated latency".
In other words, I'll start to work on your problem at 0800 local time the next work day.
Google CLI is great.
Aws is trying to get better.
GUI on the three sites is awful.
Maybe you're going off of past experiences and it's since gotten better?
From the outside, it feels like they are going to have to do something different to get some customer confidence back. Some sort of "mea culpa" with an explanation of what they are going to change.
From what I've heard in other teams (including some that own the big services you're familiar with), it can be a shitshow. Oncalls get paged 10*N times a week (typically 24/7 2-week rotation with two active oncalls, but depends), teams are leaking talent, desperately trying to hire good engineers while keeping the lights on.
Customer confidence is only an issue if there are alternative providers that don’t have problems. That isn’t the case with cloud hosting.
Customer confidence is an issue because they've had 4 notable issues in a very short timeframe.
Edit: I think you're reading something into my comments that's more than what I've said. I am curious what the AWS response to 4 major incidents in a month's time will be. That's it. At least in my circles, it is an issue for customers. I respect that it appears not to be in your circles.
Ahh. By "response", I did not mean the incident summary. I meant the overarching company response, if any. Like policies around change control, capacity additions, and so on.
All of the big cloud providers have problems, so unless there's a service with similar offerings and reach with demonstrably better reliability customer confidence is not going to matter -- people will use the least bad of the available choices.
I have multiple customers hosting on AWS, and a couple on GCP, and one on Azure. Azure by far has the most problems in my limited experience. I have servers on AWS (US East Ohio and US West Oregon) with 600+ days of uptime. None of my customers except the one on Azure have ever contacted me about an outage, in two or three years. At the same time I can read long threads like this on HN full of doom and gloom and predictions about the decline of AWS. Which customers are we talking about? I know I'm not the customer for AWS -- I don't pay for hosting.
Having moved more than a dozen companies from self-hosting and co-located hosting to AWS (or GCP, Azure if that's what they want) I can say that they are 100% happier. Relatively infrequent outages that AWS has the resources and incentive to jump all over and fix are preferable to trying to get me or some other pissed off engineer to drive in and try to figure things out.
If quality, reliability, and frequent outages drove customers away then Tesla would have withered up years ago. There are more factors in play here than what a small number of committed tech geeks (like me) think about how AWS could be doing things better.
HN has lots of these apologies that seem insincere: Here's what happened, here's how we fixed it, here's what we will do to prevent this happening again. It's a PR exercise, not something that necessarily improves my confidence.
Any complex technology at AWS scale is going to have outages and mysterious problems and glitches, all the time. That's inherent to both technology and human organizations, and the HN crowd should understand that better than lay people. These threads mainly serve as launch points for endless armchair diagnostics and proposed solutions from people who have no skin in the game.
Companies like Amazon should learn from this, but they likely won’t unless their employees unionize.
[0] https://www.theregister.com/2021/08/13/amazon_game_contracts...
During COVID everyone who didn’t need to be in the office worked from home. That was sometimes more stressful as you had lots more time in video meetings and no hallway interactions to make quick decisions on easy stuff.
For the problem reports today I am in wait and see mode. It’s not clear to me yet that AWS is the problem; could be a six-pack-and-a-backhoe kind of problem with a network provider.
If we're going into pure speculation mode I think we might also want to include that maybe their outages are staged and have factored in that the temporary bad press and paying out SLA credits is more profitable because it guides existing customers to start scaling to multiple AWS regions to offset these infrequent downtime events in 1 specific region at a time.
For example if you have hosting on US West but now to help reduce time down you duplicate your infrastructure in the EU region then I'm pretty sure the outcome here is if you were paying $6,000 a month before now you would be paying $12,000 a month (+/- any regional cost differences).
Realistically I don't believe this is the case but it's not impossible.
Not seeing any issues here, but am seeing people reporting broader internet issues at the moment. Post title seems a bit quick on the trigger to point blame.
As if they didn’t already get enough negative attention.
As a programmer I can probably guess how AWS experiences problems: programming and networking are hard problems, they get exponentially harder at scale, and there’s no known way to prevent every problem.
If AWS or some journalists gave us details we would just see 200 posts about how they could have done it better with Rust, nothing actually useful for anyone.
I presume that this rash of downtime is just the rickety structure inevitably creaking and breaking. It's probably too late now to fix things — dealing with legacy code requires patience, discipline, and understanding which the management of Amazon doesn't have. They'll just yell at people louder, and hold people "accountable" by punishing anyone where anything breaks.
Gotta love their leadership principles:
> Frugality (we'll give you a shitty laptop, shitty chair, and a shitty desk.)
> Be right a lot (are you too stupid to forecast the future?!)"
At least your fireplace will give some warmth for your relatives. Merry Christmas!
I think it's because there was a backup server that kicked in for AS16509; however, it also went down but only for 8 min.
We're not in the cloud or tech business per se, and as such our customers are not really understanding of technical issues which unfortunately means they are blaming us, and our own reputation is on the line because of AWS' shortcomings.
We did consider Google but for now OVH is the only major provider which is both reliable and secure (w.r.t. court-and-gag letters from government and intelligence agencies) as far as we're concerned. We still use Hetzner and Scaleway for some older stuff and also because of Scaleway's low prices, but it's likely we're moving everything to OVH in the future.
Now that I'm looking for it I see it everywhere.
(not entirely serious)