Amazon packages pile up after AWS outage spawns delivery havoc
detroitnews.com
detroitnews.com
[0] https://www.reddit.com/r/AmazonFlexDrivers/comments/rb3ggn/i...
What's worse, though, is that this also implies that even if an outage is localized to a particular region, data for deliveries to that region isn't multihomed, so you can't just failover flex app service to another region and keep delivering. For a service that lives and dies on being able to deliver everything faster than everyone else... that seems like a massive oversight... but also completely in-character for Amazon.
I don't doubt that many functions are difficult to failover, but a bare-bones minimum seems straightforward. For example, evidence of delivery is append-only and only needs to be globally consistent later, after a dispute.
But what then? You had a complete loss of traceability of shipments and operations, once you regain it, a junk of shipments isn't there anymore where there are supposed to be. No you have not one potential root cause for this, the outage, that could be resolved by retriggering those shipments (not loss, as you didn't deliver anything to customers) but two: the outage and some off-line shipments. In case it was just one FC, sure that would be doable. If the whole network in a complete region goes down, no way to handle that. It is much easier and safer to just stop operations until the outage is resolved, re-route orders to other regions in the meantime, and then work through the backlog. Amazon's ops are good at that, specifically because they have almost complete transparency on their material flows. Going off-line would have jeopardized that transparency, making a quick recovery after the outage all the harder.
Not every workload is of the micro size.
MR up our databases would cost around 15mil per year for a company that makes 50mil..
As far as I understand it, if you design is sane, the database/storage handles fallback and recovery.
Or maybe in other words - you need to make your service handle single machine going down without any problem - cloud or not. And there seem to be two options - it's your machine or part of a service which AWS provides to you. In second case it's on AWS to handle that and in the first case shouldn't AWS make it such that for you DC is just a parameter and they handle all virtual network and other magic?
To be super clear - I'm not arguing, just trying to learn I would love some specific examples which make the problem hard, because all these stories make me stay away from cloud which in theory is solution well worth paying extra for in a bigger company context.
It has been my experience however, that as more people use the cloud, all that "ease of use" both adds layers of complexity, and further, abstracts the backend away.
Thus, by outsourcing sysadmin tasks to AWS, no in house experise exists. People don't know how to handle correct failover, unless the platform 100% does it all for them.
Even if you create a system that is eventually consistent when availability is restored (a difficult problem all by itself, and probably needs a lot of application layer logic), it may not be worth the trouble. Warehouse workers interact with the "perfect" state of the real world, and if the computers don't have access to that, they aren't very useful.
And there's probably more than just a database. Maybe a message queue for asynchronous treatment, object storage for photos, etc.
It's certainly not an insurmountable problem, but maybe they consider the failure rate so low (it is, us-east-1 going down is like a once in a few years event) that the complexity of multi-region isn't worth it.
Apparently even some warehouses were ghost towns today without anyone sorting packages[1]
Here in Phoenix, I've had two packages supposed to be delivered today now delayed. Also, delivery on items just went from overnight delivery to 3+ day delivery as the quickest option.
[0] https://www.reddit.com/r/AmazonFlexDrivers/comments/rbf8ti/j...
[1] https://www.reddit.com/r/AmazonFlexDrivers/comments/rbbnwa/y...
Corporate words are worthless anyway.
Was it Metra or the CTA out of curiosity?
why is McDonalds dependent on AWS for their app to work? (/s)
Why shouldn’t Government IT be using the same tools regular companies use for IT?
> All service interfaces, without exception, must be designed from the ground up to be externalizable. That is to say, the team must plan and design to be able to expose the interface to developers in the outside world. No exceptions.
1: https://nordicapis.com/the-bezos-api-mandate-amazons-manifes...
If you're going to concentrate risk on AWS, it better be essentially flawless, stable, and highly-redundant.
A service that is unable to handle such a failure does not qualify as being ready for deployment. And I'm not just saying that to be a self-righteous pedant, I'm saying it because this kind of failure is statistically likely to happen at some point, so ignoring it is setting yourself up for major (and potentially very expensive) problems if you don't test for it regularly.
One major outage per year doesn't necessarily mean just a few hours of downtime, it could mean having to redeploy your entire service somewhere else, which could take several days or more if you haven't prepared for it. How many of those services can live with that?
Apparently, many can! Look at how many companies choose not to pay ransom when hit with a ransomware attack, or prefer to negotiate for days instead of buckling straight away, even if it means operations are crippled for weeks. They don't typically go bankrupt afterwards, everyone coped, life moves on.
And even more now: http://techblog.netflix.com/2013/12/active-active-for-multi-...
The companies we talk here about are most likely running a monolith app on Centos6..
But an entire region? I've never worked anywhere that decided being multi-region was a good tradeoff. At best we've replicated data to another region and had some of our management services there, so in the absolute worst case (which would need to be much worse than yesterday) we could rebuild our product's infrastructure there.
Do I agree with this approach? Not in all possible cases of course, but for my employers? Overall, yes. It mitigates the highest impact risks. Going further would have significant complexity and costs. Those companies success or failure haven't been impacted by their multi-region strategy AFAICT.
Services that truly need 100% uptime (and by that I don't mean what management say they WANT, but they are prepared to PAY) are a tiny minority, to the point that I imagine most software developers never work on one.
Even setups that i have seen to handle such cases, usually had some single point of failure somewhere.
And lest face it. Even if you do multi region on AWS, that won't protect you next time someone screws up BGP/switches/DNS or whatever and every region goes down for a while. In that case you better have failover to some other cloud vendor. Even planes or other safety critical equipment is not 100% failure proof.
"Neighbors in Tennessee city worry as Amazon packages pile up outside home"
"[The people who live in the house] have a friend who has a contract with an Amazon warehouse in China. They say whenever that friend's contract expires, she will send the packages to their house for the family to sort and then send back to Amazon for the company to sell."
It's possible to ask Amazon to remove selected items from their warehouse and send to any US address for a somewhat reasonable fee of around $0.32/item ([3]).
What happens here is the Chinese sellers send packages to the residents of the house, and then relist items under a different account, so that the counter is reset and the Amazon storage fees are low again.
That said, the residents of the house run a subpar operation.
1. https://sellercentral.amazon.com/gp/help/external/G3EDYEF6KU...
2. https://sellercentral.amazon.com/gp/help/external/help.html?...
3. https://sellercentral.amazon.com/gp/help/external/help.html?...
Actually, it's a bit more than that. Some half-decade ago, Google Cloud got stung by a coordinated attack over the holidays where attackers used stolen credit cards to build a net of GCE instances and do an attack on Tor via endpoint control. Cited by SRE in the postmortem was the relative immaturity of the cloud monitoring, logging, and "break glass" tools that SREs were accustom to in Borg... Essentially, Cloud didn't have the maturity of framework that Borg did and they felt the extra layer of abstraction complicated understanding and stopping the abuse of the service.
This report had a chilling effect internally. Whereas management had previously been encouraging people to migrate to Cloud as quickly as possible, after this incident software engineering teams and the SREs that supported them were able to push back with "Can we trust it to be as maintainable as what we already have?" and put cloud on the defensive to prove that hypothesis.
That is simply not true.
https://www.zdnet.com/article/microsoft-moves-closer-to-runn...
Office 365, Teams, Dynamics, xcloud, xbox live - they all run on Azure.
I'm sure that's changing with time.
YouTube also has its own infrastructure independent of GCP and the rest of Google
I'm pretty sure YouTube hasn't had their own infra in quite a while. When I was there ~8 years ago I think it was all integrated already. Certainly database, video processing, storage, CDN were all on core Google infra, and I'm sure the frontends were too though I don't remember looking into that explicitly.
Doing something on top of GCP rather than the normal way at Google was a huge pain. Borg tutorials and documentation were just far superior, I could get a thing running on borg in an hour from not knowing anything about borg, I spent a week trying to get something running internally on GCP but still couldn't get it right (our team wanted to see if we could run things on GCP so I was tasked with testing it, I couldn't find anyone who knew how to do it so we just gave up after I didn't make any real progress). That was the worst documented thing I've ever worked with. And even worse the internal GCP pages were probably running in california and probably weren't tested from Europe, so the page took like 2 seconds between mouse click and it responded to anything.
That was years ago though and I no longer work there, but at least back then the work to make using GCP internally seamless wasn't done. Maybe it is simpler if you run everything in it and don't need it play well with borg, but there is a reason why it isn't popular internally. And likely you wont find many engineers who left Google who recommend you will use it, since they probably didn't test it and if they did it probably was a bad experience (unless they worked on GCP).
Neither was AWS.
Source: used to work as a SWE on a flagship Azure service
so it's even more embarrassing how terrible the software is.
trolling aside - if this is true, the state of e.g: MS Teams is a travesty. Implementation of replies to messages implemented in 2021!? So many bugs, etc. it seriously damages productivity. And don't get me started on sharepoint.
I have a running list of teams bugs/flaws/inferiorities, it's currently about 30 items, will probably publish it sometime soon
1. Scale is hard and downtime is hard, HNers either recognize the struggle or appreciate their lack of experience. When AWS fails many armchair architects come out to suggest solutions but many more techies just sympathize with the Amazonians.
2. Technically, Amazon has built something impressive. It might not be what you want, or what others have, but AWS is impressive in scale and scope and even reliability. Many people share credit where due.
3. One can criticize the treatment of warehouse and delivery workers that Amazon is known for, but this has little baring on the tech workers there nor AWS generally. So AWS stories tend to be free from the social critique the company as a whole receives.
Where on Earth did you learn that?
This is the future. AWS has an outage and your car won't work.
My brother lives in Anchorage and he went hiking with his with ~week ago, and he said it was -20F (My brother was a colonel in the USAF, so not just like a weak person)
But the point is, she should have been able to just ran out to the car like other alaskans do, and not be late to work.
In sufficiently cold conditions, you need to warm the engine up before starting the car, with an electric heating element. If this was controlling one of those heating elements, the problem was presumably that the car wasn't able to start at all, not that the car couldn't be started remotely.
When I use a block heater, I let it run for a couple hours before needing the vehicle in the cold. I've got an outdoor outlet timer that I can use to start my truck's block heater around 2-3AM if I need it for something in the morning. It'll start at 0F without it, but it's exceedingly clear that it's not happy about the arrangement, so I preheat.
Whatever it is, I'm entirely unsurprised that some app or another, talking to some cloud service or another, talking to some car or another, fails silently when "impossible" things have happened. Nobody seems to consider that the cloud can fail. Even though it does, quite regularly, and reliably breaks all sorts of stuff every single time it does.
"Yes, the dumb broad didn't think to walk outside and turn the car on! I shall right this wrong with my clever internet post! Behold my intelligence!"
But for a one off occurrence? Why would you assume a car company knows what they’re doing over the person telling the story? It’s silly.
Like if I say "Humans have two feet" some midwit will come along with an article or anecdote about a person who was born without two feet.
And if I say the multi-hour outage of AWS made my friend minutes late because of her car warmer, what is going through the mind of a person who offers the solution "Did she try walking to the car and turning it on manually?" These are the kinds of people that if you met them in real life you'd quickly distance yourself from them.
It's autism.
One of the symptoms of autism is [the inability to recognize sarcasm](http://www.healthcentral.com/autism/c/1443/162610/autism-sar...) without the help of idiotic, illiterate signals like "/s". Your example has the same cause; people who are unable to understand nuance and social cues, whether in real life or in written form.
This is why I buy older cars with minimal electronics, BTW.
I'm really not sure what I'm going to do with cars in the future. The Volt is the last wave of not-always-connected-software-OTA-big-data-analysis cars, and even it had Onstar, it just tries to talk to towers that no longer exist. But there's a difference between that sort of system and what newer cars have. I suppose you can always disable the cell modem and let them have occasional wifi access to update, but... ugh.
There’s another Alaska, less cold?
I’ve read reports that even Ring doorbells don’t eh… ring because of the outage.
Of course you could. There have been remote car starters for 20+ years. They don't rely on internet nonsense to work either.
https://en.wikipedia.org/wiki/Failure_domain
Clouds aren't magic. They require a certain amount of operational confidence in order to understand that, yes, an entire region can fall out from under you at any time and it's your responsibility to detect and deploy into an unaffected region if possible.
edit: Generally, one entire region will not fail. However, core services like STS rely on us-east-1 so it's particularly susceptible to disruption.
There are also way too many successful, public cloud-native businesses running without any semblance of a Business Continuity or DR Plan among any of their teams. I wish I could name and shame some of the more egregious cases I have seen.
Empirically? No.
But seriously, instead of making every dev team consuming an AWS product hire the extra engineers required to build a system that spans multiple failure domains — which lets be honest, companies won't — why hasn't Amazon just hired the engineers required to do it for me?
> Generally, one entire region will not fail.
Laughs in global failures.
For many use cases, it’s acceptable to shrug and blame AWS for a failure. It’s harder when your high availability solution fails independently, which they almost always do more than US-East-1
And being the only one up doesn't win as many market cred points and you'd think.
For many services, it makes more sense to make it reliable than not. For other services, it makes more sense to think about the engineering of the solution in the field.
Example: McDonald’s product images on kiosks are apparently in S3 and not cached locally. Seems like a dumb idea to me, but I wouldn’t try to build a more reliable cloud storage backend to control that risk.
FYI this example hasn’t been true for a while. STS regional endpoints are generally what you should be using these days. The “global” us-east-1 endpoint still works, and may be the default for some clients, but isn’t a requirement.
At a previous job where we needed to always be up, our disaster recovery plan assumed that the us-east-1 site had been hit by a meteor (not literally, but that's how we explained it to each other to put ourselves in the mindset.)
https://aws.amazon.com/about-aws/global-infrastructure/regio...
The cli is gorgeous, the web gui is terrible.
So yes, in theory much of AWS's services are probably very reliable and distributed across AZs and regions, but in practice there's likely a whole bunch of debt where one thing gets fucked up and it cascades.
To be fair, we only run services in us-east-1, and we had zero downtime from this issue. The only issue we encountered were API calls that were failing for us, and those were for CloudWatch and CodeDeploy. We are heavily reliant on EC2, OpenSearch, and S3.
At the end of the day, building redundant systems is expensive (and may introduce whole new bugs), so they probably did the math and figured the risk of a whole region outage was less than the cost of building redundancy in some systems that are hard to make redundant.
Though, as you say, none of that's a great excuse for Amazon itself. I know Alexa devices, Ring Devices, Prime Video, imdb.com, etc, also all had issues in the first hours of this outage.
After 2017 I'll never use US-east-1 again. Hell... I should have learned that particular lesson in 2011 but it took two catastrophic failures for me to figure it out.
There are numerous threads here on HN covering the topic "why does US-east-1 suck so hard."
https://news.ycombinator.com/item?id=13756082 is just one example.
I kind of doubt that they’re getting efficient use of hardware Amazon could be using but isn’t, since if someone allocates it it’s not available for Amazon anymore.
[9:37 AM PST] We are seeing impact to multiple AWS APIs in the US-EAST-1 Region. This issue is also affecting some of our monitoring and incident response tooling, which is delaying our ability to provide updates.
You'd think if it's anything mission critical, Amazon would be following its very own Well-Architected Framework. The Reliability pillar would speak to this.
https://docs.aws.amazon.com/wellarchitected/latest/reliabili...
Although the cost to make all of Amazon commerce, logistics, and digital truly multi-region is probably an order of magnitude more than the impact of this outage.
But I suspect there are third part integrations that benefit from being on the same AZ as Amazon APIs. I’ve been having little convos all day about how I think the inter region pricing creates a perverse incentive that’s exacerbating the us-east-1 situation.
To be clear, I am not advocating for ignoring high availability setups. Just highlighting the complexity cost of it.
The key benefit of the cloud is blameshifting. It’s someone else’s problem, you just get the day off.
(of course outages can happen and are to be expected especially at Amazon's scale, it's just the bad communication and the amateurish non-redundant setup of their own core services that is shocking)
Imagine ransomware targeting Amazon and spilling into the real world logistics, how much ransom could they charge.
p.s. Your username?! I can think this must be the only site you've managed to get that handle?
And the username I use here isn't available on pretty much any other site, but was still available here in 2021.