Deming's Red Bead Experiment (2002)
maaw.info
maaw.info
Related: "A bad system beats a good person any time" does not mean "having any system, no matter how bad, is better than having even the best people and no apparent system".
> Related: "A bad system beats a good person any time" does not mean "having any system, no matter how bad, is better than having even the best people and no apparent system".
I'm a big fan of sociotechnical systems [0] where the motto is to give people complex jobs in simple organizations. Unfortunately in practice you usually see the tendency to do exactly the opposite.
Here is Dr. Deming himself performing the experiment https://www.youtube.com/watch?v=7pXu0qxtWPg
There are few authors that have taught me so much about people, motivation, systems, quality, statistics, what high-leverage effort looks like, and so on.
I first picked up a book by Deming a few years ago, and not a single day has passed that I have not had use for what he taught me through his writing.
The things he says are only becoming more and more relevant with every year. I honestly think it ought to be compulsory reading in school. The world would be a much better place that way; kinder, more efficient, and less superstitial.
Here's a concrete example from today: under my supervision is an extremely good, but also somewhat expensive contractor.
My predecessor hinted that upper management has started to become nervous that the costs will run away, and recommended that I try to control that cost by having everyone in the company ask for my permission before they used this contractor's time. I'm not a fan of blocking other people's work on my approval. Besides, it's not like I understand the nuances of each situation well enough to make a good decision.
So instead, I asked for the historic bills from that contractor and plotted them on an XmR control chart. Sure enough -- every single one in statistical control.
It's a stable system. There's no sign of increasing costs. NNo unusual amounts billed. Barring special causes, I can predict the future costs of this contractor perfectly, with zero intervention -- just from the data.
Now maybe this stable system results in too much expense, and that's a conversation about common causes worth having. But I see no reason to meddle with individual decisions. It can only make the variability worse.
It seems to be that is a demonstration of the following:
1) Take a task that can only be minimally affected by skill or effort (drawing random beads)
2) Pretend it's a task that can be affected by skill or effort, leading to natural interventions like worker incentives, praise, etc.
And together you get that 2) doesn't affect 1) at all. And maybe people feel bad about it afterwards. This illustrates the point that you can't just use worker incentives to optimize the system, you might need to change the system (e.g. suggest new tools). (Please correct me if this interpretation is way off).
I can see why this is an important point: many managers do think it's all about employee skill and effort, and don't look at the system.
But here is my criticism: I would say it's not an experiment, but instead a demonstration or even illustration. I (and many people) can easily predict the outcome of this illustration if it was just verbally described.
Also, 2) not working in situation 1) dangerously should not be read to mean 2) will never work. There can be many situations where the opposite of 1) applies: where the only thing that matters is worker skill and effort, like moving bricks from one pile to another without tools (presuming that is a necessary and irreducible task). In that case 2) can be the largest lever possible.
You're absolutely right here, of course. However, I'd argue that the "necessary and irreducible task without tools" constraint is much stricter than we see in practise anywhere in real life.
In my experience, even mundane tasks like these (I grew up on a pseudo-farm for some early years of my life, so I have seen plenty of variations of the moving-bricks task) can be optimised to the point of near elimination if performed in a system that encourages thinking.
When comparing the two neighbouring pseudo-farms where I grew up, you could easily see one operated with much higher efficiency than the other even when they had roughly the same moving-brick-like tasks to perform.
(When I first wrote my comment, I wrote just "moving bricks" but then realized that Deming was right after all: you could move bricks with tools =).)
And then, managers give praise or encouragement to do better, all based on those issues.
Note: I'm not saying some employees don't have more skills than others, only that the things out of an employee's control seem similar to the Deming red bead experiment.
(via https://news.ycombinator.com/item?id=5193898, but no comments there)
His understanding of High Output Management (a seminal book by the CEO of Intel) was so flawed that the CEO of Dropbox had to correct him.
Though Avery does hint that "output" itself is a function of values/principles execs ought to imbibe in their org:
> What executives need to do is come up with organizational values that indirectly result in the strategy they want.
> That is, if your company makes widgets and one of your values is customer satisfaction, you will probably end up with better widgets of the right sort for your existing customers. If one of your values is to be environmentally friendly, your widget factories will probably pollute less but cost more. If one of your values is to make the tools that run faster and smoother, your employees will probably make less bloatware and you'll probably hire different employees than if your values are to scale fast and capture the most customers in the shortest time.
It remains to be seen if Avery ends up building a larger company than Drew. I'm willing to bet all of $100 in my depleting bank account that they will.
> Contrary to what the post suggests, HOM does not say not that the job of an executive is to wave some kind of magic culture or "values" wand and rubber-stamp whatever emergent strategy and behavior results from that. CEOs and executives absolutely do (and must) make important decisions of all kinds, break ties, and set general direction.
Notwithstanding that Drew is a billionaire who founded a billion dollar company and Tailscale has yet to crack a large valuation, Avery basically misinterpreted Andy Groves. That misinterpretation, "CEO as a passive referee whose job is to set the culture", is only sensible to people who've never managed a large group of people.
Obviously there are differences between people, and better and worse teams. But the lesson here is about how the environment factors in, and how management can accidentally arbitrarily suppress innovation or reward luck within normal bounds of success. Or hamper themselves to failure by insisting on a broken process.
Could it be the case that “everybody goes to Jim,” and as a result, Jim gets good at helping people? Could it be that if everybody just went to Kim for 2 weeks, that her fixes might turn out to be better yet completely orthogonal method of solving the problem?
The Red Bean experiment is an antidote to rigid process and the praise/blame game as based on inspection of results. It’s a story intended for management to hear, not an absolution or dismissiveness of personal reasonability.
If you’ve hired “ready willing workers,” then looking at the results doesn’t necessarily show you who was killing it and who wasn’t.
That worker who is always “killing it” may be good at scooping up projects that always look great. That worker who is always underperforming might be maintaining essential infrastructure without which the system would fall apart.
The worker who’s killing it may be doing so by spending all their time “buttering up” a customer. The worker who appears underperforming may appear so because they spend all their time “buttering up” a customer, but someone else always lands the sale.
It’s a meditation on imperfect knowledge.
A good example of the type of mismeasurement done in non-manufacturing contexts is the ridiculously stupid burn-down chart.
Bad management can find a misuse for any tool, I don't think burn-down charts are a particularly attractive nuisance in that regard.
I've read about people who go for days after the experiment and feel bad about their subpar performance because they feel like they've let down or brought shame to their company and wonder if they couldn't have done something better.
And this is an experiment that's set up to remove any trace indivdual agency what so ever! People still beat themselves up over it.
When you experience this experiment for real, you start to forget that it's actually designed to eliminate any sort of skill.
In other words, the experiment shows how hard it is to recognise when we're judging the system and not the people in it. The experiment shows that even when you think you're seeing individual performance, it's very plausible you're not.
The content of that book isn't directly applicable to e.g. software companies but if you think a bit you can see quite a lot of analogous situations (e.g. warehoused inventory is incomplete projects or not-yet-shipped code, "monuments" could be inappropriate central test / build systems, etc).
Ward's Lean Product and Process Development is also a good take on those ideas.
So first, not all science requires an RCT. Dividing expiremental subjects into study and control groups is one way of doing science. It's not the only way.
In this case, this is a concrete demonstration of just how much variance can emerge from a "statistically neutral" process. The systemic flaws are part of the demonstration. What appear at first glance to be identical tools, inputs, and processes are in fact subtly different. The demo shows management types that their charts and graphs cannot always be relied upon to differentiate performance levels among staff. The system itself must also be scrutinized. If Bob's ad campaigns are outperforming Alice's by 20% in the first quarter, it doesn't necessarily mean Bob is a marketing genius and Alice needs a PIP.
A computer simulation would not have nearly as powerful effect on most people as a live demonstration using real beads. And the imperfections in the paddles is something that naturally arises when they're physically made, but would have to be tuned by the programmer building the simulation. Which would lead to questions about "just how did they decide what variances would come into play?"
Meanwhile another worker—previously scoring 5, and getting a merit-based raise from it—did poorly with 12. The remark: "That raise went to his head. He's getting lazy".
So yeah, the value of this is in the actual doing of it.
If you are curious (and have the stomach for some MBA-ness), this article (four parts) should be well worth your time: https://archive.is/tXJhw
This is not in their control — they’re victims of corporate policy.
Does it influence quality in complex and hard to quantify ways?
Most assuredly…
If that's his point, it seems obviously wrong. Some work, like in the experiment is dominated by system behavior. For others, system would play a much smaller role. For example, 2 people cranking code in a startup. No matter what system you apply, if they are not good programmers, nothing of value will come out.
Randomly swap in two new actors with different life experiences into the same spots to do the same work, and you'll still get shit. If in the unlikely event, you get amazing work, it's not that the people doing it were special; it's just anothe outlier in the data stream. Add in the emotional toll of working as hard as possible to succeed but never being able to meet prescribed quality levels?
A system is perfectly tuned to produce the results it does. Want different results? Change the system. That is Deming's point. We have a tendency to blame variance in a system on the human actors immediately proximal, instead of paying attention to the actual significant constraints. This is an important lesson to management types, as they are to process/system what a programmer is to a computer.
The planners cast the dice for downstream long before downstream can do anything about it, and in many corporate setups, top down works just fine, but bottom up never gets any attention.
Deming on various management topics https://deming.org/category/deming-on-management/
More resources on Deming's ideas https://deming.org/online-resources-on-w-edwards-demings-man...
[0] https://www.goodreads.com/en/book/show/34987.Four_Days_with_...
Find your strengths, find something you enjoy that utilizes those strengths, and find a career where you can stand out for the combination.
IOW, the experimenter wanted to be able to arrive at the conclusion that difference in performance was unrelated to workers and designed the experiment so it would give this result. In short, this demonstrate few things outside of a very artificially setup situation, where the workers have no say and the job is predestined to fail.
Anyone who worked anywhere knows very well that there are actually vast difference between two workers.
That's the whole point. The experiment is not that we're supposed to be surprised that the workers did not affect performance - in fact, that's the subtext of the whole thing! We know it from the start cause he explains exactly how the process works and we can all see that individuals cannot affect their output.
The point is, if we are unaware that we're in such a situation, we can still find metrics to allow us to rank workers, fire low performers, give out raises, etc. When we myopically focus on such metrics, and disregard the system that makes them worthless, we're making all our decisions on random chance, even though we have a clear process, data collection, the whole thing.
The broader points he makes, related to this experiment at least: There are individual and systemic issues that influence the outcome of a process. The actual ratio will vary depending on what kinds of processes are involved.
If the job is to be a literal screw turner on an assembly line, then there is relatively little difference between the majority of people (assuming they are generally able bodied, sighted, and have decent coordination), the system (tempo, length of shift, accessibility of the thing being screwed together, tools being used) will have a much larger impact than the individual's skill. The system of the assembly line will influence the outcome more than the individual's skill (at least above a basic threshold, a supremely uncoordinated individual could flounder even with the slowest pace of work). Switch to more skilled work and you will find, increasingly, more differences in outcome based on individual performance versus the system of the work, but even there the system matters.
Look at software development offices that still favor things like manual build processes, code versioning control, testing, and deployment over automation. They provide many opportunities for human error (even just simple miskeying of data) that can reduce everyone's effectiveness no matter how skilled. (Fortunately these kinds of places are increasingly rare, at least outside of US defense contractors.)
The experiment, then, is an artificial construct (like most classroom experiments) meant to illustrate a point by showing one extreme. This acts as a counterpoint to the more conventional wisdom that the individual, and not the system, is what actually matters for the outcome. The conventional wisdom, of course, being wrong in many circumstances since it tends to place too strong a weight on the individual performance and too weak a weight on the system.
It would be unsavory if he had said, "See, stop evaluating individuals their contribution doesn't matter." But he never did say that (in anything I read, at least), and anyone who looks at this experiment and draws that conclusion would be an idiot.
In other words, this applies even more severely to knowledge work.
More concretely: in manufacturing you can have a process that yields 9–14 % defective or whatever. The variation is relatively small; say a CV of 10 %. In knowledge work, you'll be looking at processes that generate somewhere between 0.1 and 100 defective ideas for every really good idea. This variation is enormous: 1000 % or so.
Heck I can tell you from experience that if you want to get promoted fast, new product teams are the way to go. You get to file lots of patents, architect huge new systems, and look like a rock star.
Another example: Partner teams upstream keep pushing breaking API changes, downstream teams look bad because their services are the ones having the outage. You do your due diligence, your code is defect free, well tested. Doesn't matter, you are spending half your day putting out fires caused by someone else. Meanwhile another co-worker starts on a team where their upstream services are written to be robust against bad incoming data and have APIs that maintain back compat. Your co-worker puts out buggy poorly tested code, but the upstream services are robust enough that everything keeps chugging along.
Management doesn't see any of this. They just see your team has poor performance, and this other team has great performance. Heck maybe that other team has a higher "velocity" because they can turn out features faster.
If the process was the problem (or perfect), adding a new (or replacing) a worker would have minimal impact but we know this is not true. You could have the best documentation, training, and absolutely stellar code yet one person can turn everything to shit quite quickly! The opposite is true as well: Bringing on a fantastic new worker can make your existing team look like a bunch of inefficient laggards.
Neither of these situations can be fixed by improving processes (maybe hiring processes? Though I doubt it). It'd be like having one magic blue bean in the box that--if found--can either drastically improve or degrade the final productivity by 90%. Would the optimum process improvement then be to try to eliminate magic beans entirely? Sure seems like it (i.e. hire the lowest common denominator and don't try to optimize for the 1%). That way you reduce the likelihood of taking on the "bad 1%"--even though it reduces your chances of obtaining the perfect magic bean.