From the outside, it feels like they are going to have to do something different to get some customer confidence back. Some sort of "mea culpa" with an explanation of what they are going to change.
From the outside, it feels like they are going to have to do something different to get some customer confidence back. Some sort of "mea culpa" with an explanation of what they are going to change.
From what I've heard in other teams (including some that own the big services you're familiar with), it can be a shitshow. Oncalls get paged 10*N times a week (typically 24/7 2-week rotation with two active oncalls, but depends), teams are leaking talent, desperately trying to hire good engineers while keeping the lights on.
Customer confidence is only an issue if there are alternative providers that don’t have problems. That isn’t the case with cloud hosting.
Customer confidence is an issue because they've had 4 notable issues in a very short timeframe.
Edit: I think you're reading something into my comments that's more than what I've said. I am curious what the AWS response to 4 major incidents in a month's time will be. That's it. At least in my circles, it is an issue for customers. I respect that it appears not to be in your circles.
Ahh. By "response", I did not mean the incident summary. I meant the overarching company response, if any. Like policies around change control, capacity additions, and so on.
All of the big cloud providers have problems, so unless there's a service with similar offerings and reach with demonstrably better reliability customer confidence is not going to matter -- people will use the least bad of the available choices.
I have multiple customers hosting on AWS, and a couple on GCP, and one on Azure. Azure by far has the most problems in my limited experience. I have servers on AWS (US East Ohio and US West Oregon) with 600+ days of uptime. None of my customers except the one on Azure have ever contacted me about an outage, in two or three years. At the same time I can read long threads like this on HN full of doom and gloom and predictions about the decline of AWS. Which customers are we talking about? I know I'm not the customer for AWS -- I don't pay for hosting.
Having moved more than a dozen companies from self-hosting and co-located hosting to AWS (or GCP, Azure if that's what they want) I can say that they are 100% happier. Relatively infrequent outages that AWS has the resources and incentive to jump all over and fix are preferable to trying to get me or some other pissed off engineer to drive in and try to figure things out.
If quality, reliability, and frequent outages drove customers away then Tesla would have withered up years ago. There are more factors in play here than what a small number of committed tech geeks (like me) think about how AWS could be doing things better.
HN has lots of these apologies that seem insincere: Here's what happened, here's how we fixed it, here's what we will do to prevent this happening again. It's a PR exercise, not something that necessarily improves my confidence.
Any complex technology at AWS scale is going to have outages and mysterious problems and glitches, all the time. That's inherent to both technology and human organizations, and the HN crowd should understand that better than lay people. These threads mainly serve as launch points for endless armchair diagnostics and proposed solutions from people who have no skin in the game.