It seems to me that you can have high levels of reliability, a great working culture and high levels of staffing, but if you mess with any one of these elements you will find the wheels start falling off the train.
It seems to me that you can have high levels of reliability, a great working culture and high levels of staffing, but if you mess with any one of these elements you will find the wheels start falling off the train.
The scattering of weird microservices I think is the same reason the Amazon detail & home page never get a full ground-up update. It's hard to do large scale data experiments to demonstrate the benefit, and it's too many teams with too much at stake (too many cooks)
Are they testing for the UI's that confuse customers the most? \s
The Expedia example really comes to mind too of micro-optimizations and everyone working towards their own org's goal can all make sense individually, but really fail when taken as a whole: https://www.qualtrics.com/blog/upstream-thinking-saved-exped...
Expedia had an example of this: - the sales team wanted sales - the phone support team wanted to turn over calls quickly -
If there were a way to measure confusion, it would be improved and
Can I A/B test whether there will be a 10% increase in sales if I lower prices on Ec2 instances by 5% - YES! Can I A/B test that the presence of 30 service options is the tipping point of spending 1 hour to get something done vs wanting to hire an "AWS"
I believe it was an Amazon VP that I was talking who pointed out that the homepage and detail page had not changed very much over time at all (lots of evolution, but never revolution, never fully redone, never will be fully redone).
This was pertinent because we were in an org that ran offshoot sites, similar in a lot of ways, but less hands in the pot & we did have liberty to do a full rewrite of our detail pages & homepages. In the conversation, the VP of that org was highlighting that flexibility and went into small detail why that was not the case for the big mother-ship retail pages.
https://aws.amazon.com/executive-insights/content/how-do-you...
AWS does have differences for sure - but the "everything is only decided based on numbers unless you are Jeff Bezos" is universally true.
Being that rigorous about data-driven-decisions is really quite powerful. The S-team are 1000% of this mind-set. I'm paraphrasing, the saying at Amazon that they tell new recruits and pride themselves over is that there are only three correct answers: "(1) Yes, and here is the data why. (2) No, and here is the data why. (3) I don't know, and I'll have the data shortly"
You can make still great decisions based on your data without doing the thing where you ship the same feature twice and see which one works better.
By way of clarification, I rarely observed the same feature being shipped twice. It's generally more the 'A' test was the existing functionality and 'B' was whatever change you wanted to make.
The A/B testing platform at Amazon received a lot of investment - it's powerful and so it's used for a lot more than mere A/B testing. It's also a feature toggle platform, lots of things are shipped under that framework so they can be "turned on & off". The metrics collection was powerful and highly integrated. Even if there is no actual comparison in parallel, the metrics that are generated were worth a lot (and often needed, as everyone's goals are all data based, so you want data to show you are hitting those goals. Can't just say we launched 5 features, and can't say we are just still here, it's data or nothing at Amazon)
Most importantly, the "great decisions" is interesting. The power of gathering data very quickly and often is you see when decisions were not great. Amazon was excellent at this, things that made more money were given support, things that lost money were quickly terminated and re-organized. The data provides light on the decisions, it's virtually a BizDev super-power to be able to pivot away from losers so quickly like that.
I actually think that the "pizza team" model in Amazon can help mitigate this, as each individual team operates as its own entity that is discouraged from relying on other teams for uptime... but this is only effective to a point.
What it also means is that they can use younger and cheaper workers. They can also outsource their work. Thus my concerns about reliability.
It seems like a house of cards. I truly wonder for how long they can keep this up. Given our reliance on cloud computing and their prominent position within it, it may become a very rocky ride at some point for businesses who have transitioned their operational infrastructure to AWS. I think a lot of people treat AWS as an infinite resource that will never run out and will be completely reliable forever. I think, sadly, many organisations may get a shock some time into the near future.
Maintainability and readability are best solved through brute force hiring instead of better practices.