Foundation · 13 min
The best worker, and why they are not
After this you can
- explain why ranking people on the output of a stable process measures luck rather than performance.
Assumes you have done Signal or noise: reading a number that moves.
The problem
A contact centre ranked its twelve agents every month on first-call resolution and published the table. The bottom two went on a performance plan; the top one got a voucher. Over fourteen months, nine of the twelve had been in the bottom two at least once, and four of them had been top. Nobody's behaviour had changed. Every agent took the calls that arrived. The table was a ranking of which calls each of them happened to get.
The idea
The previous lesson was about reacting to a process that moved. This one is about the more expensive version: reacting to people who differ.
When several people do the same work in the same system, their results differ. That is guaranteed — identical performers produce different numbers, because the work that reaches them differs. The question is whether the difference between two people is bigger than the difference the system produces on its own.
Almost always, it is not. And almost always, somebody ranks them anyway.
What a ranking of identical things looks like
It looks exactly like a ranking of different things. Somebody is top. Somebody is bottom. The gap between them is real, arithmetically. You can put it in a table and it will be correct.
The tell is what happens next period. A ranking of genuinely different performers is fairly stable: the good ones tend to stay near the top. A ranking of identical performers reshuffles — last month's best is regularly this month's worst, and over a year most people will have visited both ends.
That is not a flaw in the measurement. It is what a ranking of identical things must do.
Deming's demonstration
Deming used to run this with a bowl of 4,000 beads, 20% of them red, and six volunteers he called "willing workers". Each drew 50 beads with a paddle. Red beads were defects.
He would praise whoever drew fewest, put whoever drew most on notice, set a target, and run it again. Then again. He was, deliberately, an entirely reasonable manager doing entirely normal things.
Nobody could influence the result. The paddle takes what it takes. Every difference between the workers, and every difference between days, was the bowl.
The point is not that ranking is unfair. It is that ranking a common-cause system produces information about nothing — and then acts on it, which does real damage:
- The person praised learns that something they did was right. It was not, so they cannot repeat it.
- The person on notice learns that effort does not change the outcome, which is true and demoralising.
- Everyone learns that the measurement is arbitrary, and behaves accordingly.
Targets, and the third option
Set a target below what the system delivers and you have created something that cannot be met by effort. Two things can still meet it: luck, and changing what gets recorded.
The second is not a moral failure of the person. It is the predictable output of an impossible target, and it arrives with a slow degradation of the data everyone else depends on. Once the numbers are being managed rather than measured, no improvement work is possible at all — you have lost the instrument.
What to do instead
Plot the results as one process, on one chart. If everyone falls inside the same limits, there is no difference to explain, and the honest report is "these people are performing the same". That is a real, useful finding.
If somebody genuinely is outside the limits, that is worth investigating — and the first question is still what is different about their work, not about them. Different call mix, different shift, different equipment.
Improve the system, because it is the only thing that will move the number. In the bowl, the only way to get fewer red beads is fewer red beads in the bowl. No amount of attention to the workers changes a 20% bowl.
Worked example
Six workers, six days, a bowl that is 20% red, and a paddle of 50. Here is the league table the foreman would publish.
| Rank | Worker | Total red beads |
|---|---|---|
| 1 | Worker 1 | 50 |
| 2 | Worker 3 | 54 |
| 3 | Worker 5 | 56 |
| 4 | Worker 2 | 60 |
| 5 | Worker 4 | 62 |
| 6 | Worker 6 | 62 |
Worker 1 against worker 6 is 50 against 62 — a 24% difference. On this table, worker 1 is clearly the better operator. Any manager would read it that way, and the arithmetic is correct.
Now plot every draw as one process. An np chart over all 36 draws gives a centre line of 9.56 and limits of 1.22 to 17.9 — the tool rounds to two decimals and trims a trailing zero, and the exact values are 9.5556 and 1.2150 to 17.8961.
Every single draw falls inside them. The lowest was 6 and the highest 16. Not one point is outside the limits, which means there is no difference between these workers, and no difference between the days, that needs explaining. It is one process, and it produced all 36 numbers.
The reshuffling. Over the six days, the leader changed on 5 of the 5 day-to-day comparisons — every single day a different person was top. Three of the six workers led at some point, and on three days the previous day's best became the worst or the reverse.
That is what the fourteen months at the contact centre looked like, compressed.
The target. The foreman set a target of 3 defects. The bowl is 20% red and the paddle takes 50, so the process averages 10. Across all 36 draws, not one met the target. There was never a possible day on which anybody could have.
What the chart says to do. Nothing about the workers. The centre line is 9.56 because the bowl is 20% red; the only intervention that changes the output is changing the bowl. Every minute spent on the ranking was a minute not spent on the only thing that would work.
Your turn
The experiment opens with six workers over six days.
- Look at the league table first. It is genuinely persuasive — somebody is best, somebody is worst, and the gap is large. Sit with that for a moment before scrolling.
- Now read the chart underneath. Every draw inside the limits, and the tool says so in words.
- Press Run it again. The table completely reorders and the chart says exactly the same thing. Do it three or four times and watch how convincing each new table is.
- Set the target to 3 and check how many of the 36 draws met it. Then set it to 10 — the process average. Rather more than half now meet it, and the exact count moves every time you reseed, which is the lesson twice over. For this bowl the true figure is 58.4%: the tool counts a draw as meeting a target of 10 when it lands ON 10 as well as below, and 10 is the single most likely result, so including it lifts the pass rate from 44.3% to 58.4%. A target set at the average is met by most people, most of the time, because of where the counting boundary falls — not because of anything they did.
- Raise days to 15 and look at how many of the workers have led at least once. Given long enough, nearly everybody gets a turn at both ends.
- Now change the bowl to 5% red. Every worker improves immediately, and none of them did anything. That is the only intervention that ever worked.
The Red Bead Experiment
practiceA bowl of 4,000 beads, 20% of them red. Each willing worker draws 50 with a paddle. Red beads are defects. Nobody can influence the result — the paddle takes what it takes.
| Rank | Worker | Day 1 | Day 2 | Day 3 | Day 4 | Day 5 | Day 6 | Total |
|---|---|---|---|---|---|---|---|---|
| 1 | Worker 1best | 7 | 11 | 8 | 8 | 6 | 10 | 50 |
| 2 | Worker 3 | 8 | 7 | 11 | 6 | 13 | 9 | 54 |
| 3 | Worker 5 | 13 | 9 | 10 | 7 | 8 | 9 | 56 |
| 4 | Worker 2 | 11 | 6 | 11 | 16 | 6 | 10 | 60 |
| 5 | Worker 4 | 12 | 10 | 9 | 8 | 13 | 10 | 62 |
| 6 | Worker 6worst | 9 | 12 | 10 | 8 | 8 | 15 | 62 |
Worker 1 drew 50 red beads and worker 6 drew 62 — a difference of 12. On this table one of them is plainly better than the other. Now look at the chart.
Show the data behind this chart
| # | red beads | LCL | UCL | Beyond limits |
|---|---|---|---|---|
| 1 | 7 | 1.22 | 17.9 | no |
| 2 | 11 | 1.22 | 17.9 | no |
| 3 | 8 | 1.22 | 17.9 | no |
| 4 | 12 | 1.22 | 17.9 | no |
| 5 | 13 | 1.22 | 17.9 | no |
| 6 | 9 | 1.22 | 17.9 | no |
| 7 | 11 | 1.22 | 17.9 | no |
| 8 | 6 | 1.22 | 17.9 | no |
| 9 | 7 | 1.22 | 17.9 | no |
| 10 | 10 | 1.22 | 17.9 | no |
| 11 | 9 | 1.22 | 17.9 | no |
| 12 | 12 | 1.22 | 17.9 | no |
| 13 | 8 | 1.22 | 17.9 | no |
| 14 | 11 | 1.22 | 17.9 | no |
| 15 | 11 | 1.22 | 17.9 | no |
| 16 | 9 | 1.22 | 17.9 | no |
| 17 | 10 | 1.22 | 17.9 | no |
| 18 | 10 | 1.22 | 17.9 | no |
| 19 | 8 | 1.22 | 17.9 | no |
| 20 | 16 | 1.22 | 17.9 | no |
| 21 | 6 | 1.22 | 17.9 | no |
| 22 | 8 | 1.22 | 17.9 | no |
| 23 | 7 | 1.22 | 17.9 | no |
| 24 | 8 | 1.22 | 17.9 | no |
| 25 | 6 | 1.22 | 17.9 | no |
| 26 | 6 | 1.22 | 17.9 | no |
| 27 | 13 | 1.22 | 17.9 | no |
| 28 | 13 | 1.22 | 17.9 | no |
| 29 | 8 | 1.22 | 17.9 | no |
| 30 | 8 | 1.22 | 17.9 | no |
| 31 | 10 | 1.22 | 17.9 | no |
| 32 | 10 | 1.22 | 17.9 | no |
| 33 | 9 | 1.22 | 17.9 | no |
| 34 | 10 | 1.22 | 17.9 | no |
| 35 | 9 | 1.22 | 17.9 | no |
| 36 | 15 | 1.22 | 17.9 | no |
Expected per paddle
10.0
Expected σ
2.81
Outside the limits
0
Leader changed
5/5
The findingEvery single draw is inside the control limits. There is no difference between these workers to explain, and no day that needs accounting for. The bowl produced all of it.
The ranking reshuffles. The leader changed on 5 of 5 days, and 3 of the 6 workers topped the table at least once. On 3 days the previous day's best became the worst, or the reverse.
The target0 of 36 draws met the target of 3. The bowl is 20% red, so the process averages 10.0 defects per paddle — a target below that cannot be met by effort. It can be met by luck, or by changing what gets recorded.
How this is calculated
Each paddle takes 50 beads out of the bowl without replacement, so the draw is hypergeometric rather than binomial: at each bead the chance of red is (reds remaining ÷ beads remaining). That is the physical experiment. The finite-population correction is only about 0.6% on the standard deviation here, and modelling it as binomial would be modelling a different apparatus while claiming to model this one.
Expected red per paddle = n·p = 50 × 0.20. Standard deviation = √(n·p·(1−p)·(N−n)/(N−1)).
The control limits are an np chart computed from the draws themselves, by the same function the Control Chart tool uses — not from the known bowl proportion. A chart that used the answer would be assuming what it is meant to demonstrate.
Source: Deming, W.E. (1986), Out of the Crisis, Ch. 11; Deming (1994), The New Economics, Ch. 7.
Check yourself
No hints. Wrong answers are explained, not softened.
Twelve agents are ranked monthly. Over a year, nine of them have been in the bottom two at least once. What does that pattern indicate?
A process averages 10 defects per batch. A target of 3 is set and bonuses attached. What are the three possible outcomes?
On a chart of all operators, one person's results sit consistently above the upper control limit. What is the first question?
Worth remembering
How can you tell a ranking is measuring noise?
It reshuffles. A ranking of genuinely different performers is fairly stable; one of identical performers sends most people to both ends within a year.
What are the three ways to meet a target set below the process average?
Luck, changing what gets recorded, or changing the system. Effort is not one of them — and the second destroys the measurement everyone depends on.
Someone is genuinely outside the control limits. What is the first question?
What is different about their WORK — the mix, the shift, the equipment, the training. The system explanation is both more likely and more fixable than the person one.
Can you do this now?
Rate yourself honestly. We compare your rating with how you actually answered — the gap is more useful than either number alone.
I can explain why ranking people on the output of a stable process measures luck rather than performance.
Your rating is recorded alongside your drill results. Neither alone marks the competency as met.