Amazon search was down
techcrunch.com
techcrunch.com
Is it a strange indicator that this feels more like lockin to me, than loyalty? I have no idea what that means.. or how to change the conceptual impression.
My hunch would be that your subconcious is telling you loyalty towards a huge faceless corporation feels weird. If it was a local store with people you could connect your positive experience to, my guess is you'd have no issue calling it loyalty, even if that store was part of an equally big corporation.
With that said, it's hard to say if that's because i hear good things, or because it's a physical store with people i see.
Compare that to Amazon, and i rarely hear good things about their staff treatment. Despite having a good experience with Amazon. .. to be clear, i just don't hear good things.. not implying that i hear lots of bad.
This is in stark contrast to many other faceless BigCo's that I've stopped using entirely, or only use out of necessity but loathe every interaction with. (eBay, PayPal, Comcast, Verizon, etc.)
Impulse purchases at the click of a mouse have really changed the way I personally consume. Anyone aware of any studies on the impact of instant feedback?
It's too bad they haven't expanded faster - Walmart just barely rolled out their curb side pickup service where I live - if Amazon had made it here with same day delivery even just a few months ago we'd be customers for life, but as it is we'll probably stick with Walmart for groceries even when Amazon does eventually get here. (better the devil you know, etc.)
I'm sure if you totaled up my Amazon purchases and my Steam library it's cost me thousands.
It just happened to me, I was trying to search a book and want to buy it, but since I can't search, I didn't. And now I don't want to buy it anymore
(1 - (365 * 24 - 4) / (365 * 24)) = 0.000456621...
0.000456621... * 100,000,000,000 = $45,662,100.4566...As another poster mentioned, even with the post-outage spike in sales from folks coming back to get what they wanted an hour ago, they lose those impulse sales forever.
But some won't. They'll either be buying on a whim (and the mood passes), or buy from a competitor. Those are lost sales that aren't coming back.
There are people around who never buy on a whim, but there are also a lot of people that buy stuff because the thought crosses their mind, forget about it, then remember when it turns up several days later. These aren't always things you "need" strictly speaking.
As to competition I'd name Jet.com, Target, and Walmart as being major ones. I buy a LOT of stuff from Target these days using the Red Card (effectively 5% pays for the tax, and free delivery on all purchases).
Behavior over 100ms can't be compared to behavior over hours. It's possible that it's true, but it's unlikely.
On the other hand, you can model some buying as periodic, not need-based. I go to the grocery store when I need food, but I go to J Crew once every six months. If I'm traveling this June, I'll delay my trip until July, so while the revenue isn't "lost", it's pushed into the future. My next trip is then in January, not December, and at this point, J Crew has lost one month of my sales forever.
There's a similar effect for many items at Amazon. I only spend so much time shopping, so if I can't do it now, there's some opportunity they miss out on in the future. The aggregate ends up looking like revenue that's lost forever.
And don't forget the CPU throttling, noisy neighbors, IOPS and what not.
For people who use it as a VPS provider, they are going to be hard hit by other tenants who are applying the intended purpose.
When things don't work, AWS doesn't tell you a reason why it didn't work. Instead they teach you lessons about how to architect your applications for failure by saying:
The EC2 team also recommend that you architect for failure using the white paper linked below:
https://d0.awsstatic.com/whitepapers/AWS_Cloud_Best_Practices.pdf
See my earlier comment as well: https://news.ycombinator.com/item?id=11822298 The EC2 has responded, and informed me that they do not provide analysis of individual issues in AWS infrastructure.
They referred to the AWS SLA that states that "AWS will use commercially reasonable efforts to make Amazon EC2 and Amazon EBS each available with a Monthly Uptime Percentage (defined below) of at least 99.95%".
The AWS SLA is available here:
https://aws.amazon.com/ec2/sla/
(The above is a part of the response I received when I attempted to ask AWS about sporadic reboots and outages on one of my instances.) Thanks for that information. It really help us understand the issue faced. As you know changing a A record for a domain can take time to replicate through all the Root DNS servers and then onto the none authoritative servers from there on. When you are editing your DNS Zone file it is highly related to the TTL settings for quicker updates on changes to them.
However it really isn't uncommon to see full DNS replication when making changes to an A record.
..
Also for even more control and perhaps faster DNS resolving times within AWS look into our Route53 service.
https://aws.amazon.com/route53/
I hope this has addressed your questions and please feel free to ask if there is anything else.
Tools like http://mxtoolbox.com/dnspropagation.aspx and http://www.viewdns.info/propagation/ indicated that the change had been picked up by every listed server but for AWS.Solution proposed by AWS: Use Route53.
There are a ton of these teams. AFT (several hundred devs), for example, is practically on their 10th year of a 3 year mandate to get off of their monolithic Oracle database backends, a constraint which forces them to run in legacy datacenters. You'll find similar issues with plenty of Supply Chain, Operations, and Transportation teams as well. The primary data warehouse clusters (non-redshift) and management interfaces are also in legacy data centers, last I recall.
I imagine a room full of Amazon engineers all standing around freaking out and possibly yelling at interns.
There are still areas of Canada (Southhampton Island) and Central America (Panama IIRC) that are on EST year round.
The first and most important principle that every call leader promotes is "stay calm". It would be natural for folks to freak out and get stressed when the clock is ticking on an outage. There's an innate human sense that our emotional state has to match the urgency of a problem. But this would only lead to a kind frenetic energy that isn't helpful.
Instead, maintaining a collected composure quickly cools everything down so that actions can be thought through and communicated more clearly. Many people are surprised that when they join a call, especially with an experienced team, that it is so calm. If you listen to the audio of Michael Schumacher as an F1 driver, or to fighter pilots performing complex maneuvers, it can also be striking just how calm they are. It's very effective and a great lesson for teams under pressure.
In that situation, and especially at scale, it's very useful to keep everyone on the same page about what actions are happening when, and to help folks out by showing them how to find things which may have changed (keep in mind the change may have been in a totally different system than the one that ultimately breaks). For example: it would slow things down if someone tried to flip data centers at the same time that someone else was rolling back software.
If on a call it becomes clear that a team needs to make some kind of non-rollback change, like a code change, or push a new config, and thankfully that's very rare, we'll generally recommend they break off and work on it in isolation, while nominating a point person to liaise with everyone else. We do use chat too and that can be more effective as more people join.
For postmortems we try hard to do in writing first, and then an in person review. It's very helpful to be able to re-assure and guide each other. I can back up the "Blame the system, not the engineer" guidance others have written about here too.
In a microservices or SOA architecture, services can layer on top of other services, that call more services, that put messages in a queue for handling by yet another service, etc., belonging to different teams. Unwinding this to identify the problem is something that's much faster with a phone call. "Connecting the dots", especially across teams. A technical discussion within a team is likely to take place by chat.
I've participated in a number of tech conference calls over the years, and I've always been impressed with the professionalism: the discussion is calm and focused on identifying the problem and recovering. I think folks strive for the "cool radio voice", the Houston Center Voice like was described in the SR-71 Ultimate Ground Speed Check story: http://oppositelock.kinja.com/favorite-sr-71-story-107912704...
> Now the thing to understand about Center controllers, was that whether they were talking to a rookie pilot in a Cessna, or to Air Force One, they always spoke in the exact same, calm, deep, professional, tone that made one feel important. I referred to it as the " Houston Center voice." I have always felt that after years of seeing documentaries on this country's space program and listening to the calm and distinct voice of the Houston controllers, that all other controllers since then wanted to sound like that, and that they basically did. And it didn't matter what sector of the country we would be flying in, it always seemed like the same guy was talking. Over the years that tone of voice had become somewhat of a comforting sound to pilots everywhere. Conversely, over the years, pilots always wanted to ensure that, when transmitting, they sounded like Chuck Yeager, or at least like John Wayne. Better to die than sound bad on the radios.
Godspeed, Amazon site reliability team...
Last quarter amazon earned $20.58bn from product sales.
So 4 hours of that is $38,000,000.
Still, some people who do go somewhere else, or are less likely to try Amazon in the future. I'd be surprised if this wasn't at least a multi-million dollar mistake.
If I go to Amazon.com I just go to the search bar and refine from there. So I guess I use their "navigation" in the form of their groupings of features on the left side.
Any calculation based on the hours down needs to be reduced to account for people finding products by other means (already saved for later in their basket, direct links, searches via external search engines).
I'd say that number is probably a ceiling on their losses, and that the actual number should be much lower. I'm super curious if non-tech news organizations will pick up this story, in which case the biggest hit could be from the PR. That being said, it doesn't seem like such a bad time to go down.
Mistakes happen. Sometimes they are really expensive.
I've got a friend who's a manager at the Amazon Game Studio. That part of the company at least doesn't do stack ranking in the way that it was described in the NYT article. There may be other teams (or even most of the company) that do, though. I think that teams in different locations end up developing their own internal culture, to a certain degree.
When I read it I remember thinking how true it was. That person would be very unlikely to commit the same error, and might even work to prevent others from making that error.
So the system failed, not a particular person, and the solution is to dive deep into why the system failed, and improve it.
Obviously there are exceptions for willful extreme negligence or policy violations.
As long as you can add to your cult karma, you can do anything at Amazon!
I have several friends that work at, or worked at AWS within Amazon. When something goes wrong, I hear it's a blame game. I hear that it's a witch hunt and the person that's responsible gets held accountable, even if they didn't cause it. Example: Employee is in charge of project, project hires subcontractors, subcontractors screw up, Employee is fired.
Sounds unfair. Sounds unreasonable. Happened.
Now, that's not to say there's not a lot of pressure. People are held accountable. But they are held accountable for fixing the underlying issue and making sure the same mistakes aren't made again. I've never heard of someone getting fired for making a mistake. People get fired when the same expensive mistake is made repeatedly by the same person. (And really, based on my experience, that might not even be enough. I think in order to get fired you probably would have to fail to understand the nature of the mistake when everyone else in the room understands.)
But of course the real reason is that Amazon's work environment is basically chock full of shitty middle-managers, many who are adopted from Microsoft or salvaged somehow from better tech companies in the valley.
They even have these leadership principles that are basically a bunch of no-brainers that they force down the employees throats at every opportunity. They use them as a criteria for performance reviews, even before actual hard metrics. Shitty middle managers will keep stressing the leadership principles because (1) that's what they are told to do (2) they need to rate their subordinates on them (3) because these principles are qualitative, the only insurance they have is that they spewed them enough to get through.
Amazon, to sum it up, is just, shitty employees from other tech companies. And dumb/average intelligence follow-the-pack type drones - many who can hardly speak English and are on work visas..and they put many of these terrible speakers as managers!! (Amazon loves cheap labor, and will use it for ANY position possible)
When your leadership principles are sold as prolific and something every employee "should study" (Bezos words') you really can't expect to hire intelligent folk who don't want no-brainer basic business and character traits hamfisted to them as if it is the new work of Plato himself.
People get fired for leaks and being intentionally malicious, not mistakes. Even expensive ones. [besides, a big outage is usually caused by multiple smaller issues if your systems are sufficiently reliable. Also, read the SRE book about "blameless postmortems"]
Although this was longer and more severe.
Would love to see something come out about how it affected their users.
(Actually what is Amazon search used for?)
The search feature being broken translates to sales being close to zero for that time period.
Wow, turns out that's actually a thing.
You made my day! :-)
tl;dr, Amazon's product search engine actually is handled mainly at A9. They seem to attract decent talent.