Specifically, they set a list of 'outcome metrics' that they wanted to optimize. Using these WBRs they ensures that the metrics by which their middle managers are measured remain effective for actually optimizing these outcome metrics. However, at some point these outcome metrics stopped being good targets. They have led amazon to exploit their workers, and have negative impact on outside businesses and the environment to the cost of massive political pressure. Moreover, these targets have led amazon to often sell counterfeit products, as well as incredibly low-quality budget alternatives. This is slowly causing people to distrust the amazon store for many things.
In other words, whatever outcome metrics they set as their target, have seized being a good metric for how well amazon is doing.
Or maybe the problems I identify are well considered trade-offs, and they are aware this is what their targets are causing...
Unless your concerns lead to >$500B loss, they are merely tradeoffs in the overall optimization.
1. The leadership team should trust each other's data. Over a period of time, that has not been true. 2. The group reviewing the data depends on everyone being there long enough to understand and get the feel. That is also hard to achieve and imho, is no longer true. 3. The metrics tend to ignore the cost to the workers. After all, everything looks like numbers and charts.
For me, WBR is a mechanism, but Amazon's focus on just the metrics has hurt its workers, employees, sellers and partners a lot.
The following is purely my opinion:
The constant churn of employees and directors at the lower levels meant the data is mostly massaged or explained away.
The really good old-timers would focus on exceptions and single anomalies - much to the frustration of poor engineers and managers. But they intuitively knew that those small issues can snowball and catching them early helped mitigating the issue. But as the systems grew more complex and the number of teams expanded massively, the feel for the pulse became more mechanical and remote. New managers, directors and VPs did not come from a similar culture and the constant politics sidelined the more experienced employees who quit. Once that knowledge is lost, it takes multiple years to get it back and I do not think Amazon got that mojo anymore.
It felt more mechanical and ritualistic, rather than genuinely trying to uncover issues (which is when I quit). It is Goodhart's law at the next level.
This is the whole message of Goodhart's Law. Once you set a target, people will naturally aim for it. If you have to ignore the target or even not have one then you are saying Goodhart's Law is useful, which counters his narrative. The point of the measure should be to detect anomalies in the system, not drive targets or goals.
If you measure widgets per week, you can look back at the end of each week and ask the questions: * Why did we produce 10% more widgets this week than our 6 week rolling average? * Why did we produce 5% less widgets this week than our 6 week rolling average? * Is there anything we can easily do to increase the amount of widgets we produce next week?
None of these questions have a target. They are process oriented. We aren't saying what the rolling average must be, we are using it to detect variance. Or, we are investigating to see if we can produce positive variance. What we try might fail and the variance is negative but then we revert our process. Or it is positive and we have improved our process. Rinse and repeat