Is our best performer actually better, or just luckier?
Six workers, one bowl, and a league table nobody can influence. Watch a ranking form, then watch the control chart show there was never a difference to rank.
The Red Bead Experiment · foundations · Foundation · free, no account needed
Use this when
- Somebody is being ranked on an outcome they do not control
- A target has been set below what the process can deliver
- A league table is about to decide pay, promotion or a supplier's contract
The Red Bead Experiment
sandboxA bowl of 4,000 beads, 20% of them red. Each willing worker draws 50 with a paddle. Red beads are defects. Nobody can influence the result — the paddle takes what it takes.
| Rank | Worker | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 | Day 6 | Total |
|---|---|---|---|---|---|---|---|---|
| 1 | Worker 1best | 7 | 11 | 8 | 8 | 6 | 10 | 50 |
| 2 | Worker 3 | 8 | 7 | 11 | 6 | 13 | 9 | 54 |
| 3 | Worker 5 | 13 | 9 | 10 | 7 | 8 | 9 | 56 |
| 4 | Worker 2 | 11 | 6 | 11 | 16 | 6 | 10 | 60 |
| 5 | Worker 4 | 12 | 10 | 9 | 8 | 13 | 10 | 62 |
| 6 | Worker 6worst | 9 | 12 | 10 | 8 | 8 | 15 | 62 |
Worker 1 drew 50 red beads and worker 6 drew 62 — a difference of 12. On this table one of them is plainly better than the other. Now look at the chart.
Show the data behind this chart
| # | red beads | LCL | UCL | Beyond limits |
|---|---|---|---|---|
| 1 | 7 | 1.22 | 17.9 | no |
| 2 | 11 | 1.22 | 17.9 | no |
| 3 | 8 | 1.22 | 17.9 | no |
| 4 | 12 | 1.22 | 17.9 | no |
| 5 | 13 | 1.22 | 17.9 | no |
| 6 | 9 | 1.22 | 17.9 | no |
| 7 | 11 | 1.22 | 17.9 | no |
| 8 | 6 | 1.22 | 17.9 | no |
| 9 | 7 | 1.22 | 17.9 | no |
| 10 | 10 | 1.22 | 17.9 | no |
| 11 | 9 | 1.22 | 17.9 | no |
| 12 | 12 | 1.22 | 17.9 | no |
| 13 | 8 | 1.22 | 17.9 | no |
| 14 | 11 | 1.22 | 17.9 | no |
| 15 | 11 | 1.22 | 17.9 | no |
| 16 | 9 | 1.22 | 17.9 | no |
| 17 | 10 | 1.22 | 17.9 | no |
| 18 | 10 | 1.22 | 17.9 | no |
| 19 | 8 | 1.22 | 17.9 | no |
| 20 | 16 | 1.22 | 17.9 | no |
| 21 | 6 | 1.22 | 17.9 | no |
| 22 | 8 | 1.22 | 17.9 | no |
| 23 | 7 | 1.22 | 17.9 | no |
| 24 | 8 | 1.22 | 17.9 | no |
| 25 | 6 | 1.22 | 17.9 | no |
| 26 | 6 | 1.22 | 17.9 | no |
| 27 | 13 | 1.22 | 17.9 | no |
| 28 | 13 | 1.22 | 17.9 | no |
| 29 | 8 | 1.22 | 17.9 | no |
| 30 | 8 | 1.22 | 17.9 | no |
| 31 | 10 | 1.22 | 17.9 | no |
| 32 | 10 | 1.22 | 17.9 | no |
| 33 | 9 | 1.22 | 17.9 | no |
| 34 | 10 | 1.22 | 17.9 | no |
| 35 | 9 | 1.22 | 17.9 | no |
| 36 | 15 | 1.22 | 17.9 | no |
Expected per paddle
10.0
Expected σ
2.81
Outside the limits
0
Leader changed
5/5
The findingEvery single draw is inside the control limits. There is no difference between these workers to explain, and no day that needs accounting for. The bowl produced all of it.
The ranking reshuffles. The leader changed on 5 of 5 days, and 3 of the 6 workers topped the table at least once. On 3 days the previous day's best became the worst, or the reverse.
The target0 of 36 draws met the target of 3. The bowl is 20% red, so the process averages 10.0 defects per paddle — a target below that cannot be met by effort. It can be met by luck, or by changing what gets recorded.
How this is calculated
Each paddle takes 50 beads out of the bowl without replacement, so the draw is hypergeometric rather than binomial: at each bead the chance of red is (reds remaining ÷ beads remaining). That is the physical experiment. The finite-population correction is only about 0.6% on the standard deviation here, and modelling it as binomial would be modelling a different apparatus while claiming to model this one.
Expected red per paddle = n·p = 50 × 0.20. Standard deviation = √(n·p·(1−p)·(N−n)/(N−1)).
The control limits are an np chart computed from the draws themselves, by the same function the Control Chart tool uses — not from the known bowl proportion. A chart that used the answer would be assuming what it is meant to demonstrate.
Source: Deming, W.E. (1986), Out of the Crisis, Ch. 11; Deming (1994), The New Economics, Ch. 7.
How this is calculated
Each paddle draws 50 beads from a bowl of 4,000 WITHOUT replacement, so the draw is hypergeometric: at each bead the chance of red is reds remaining over beads remaining. Expected red per paddle = n p; standard deviation = sqrt(n p (1-p) (N-n)/(N-1)), including the finite-population correction. The control limits are an np chart computed from the draws themselves by the same function the Control Chart tool uses, never from the known bowl proportion — a chart built on the answer would assume what it is meant to demonstrate. With a bowl at 0% or 100% red no chart is produced, because the binomial standard error is zero and the limits would collapse onto the centre line.
Source: Deming, W.E. (1986), Out of the Crisis, Ch. 11; Deming, W.E. (1994), The New Economics, Ch. 7 — the red bead experiment and what it says about ranking people.
Learn the method
- The best worker, and why they are notFoundation · 13 min · free