Master Black Belt · 15 min
The study you can afford, and what it can actually see
After this you can
- state the four quantities a study design fixes — the difference worth detecting, the sample size, the significance level and the power — and solve for any one of them given the other three.
- compute the sample size a comparison needs from the smallest difference worth acting on, and say why that difference is a business decision rather than a statistical one.
- say what a study whose sample is already fixed by circumstance can and cannot detect, and decide on that basis whether it is worth running at all.
- explain why a non-significant result from an underpowered study is uninformative rather than evidence that there is no effect.
- show that power computed from the effect already observed adds nothing to the p-value, and name the question to ask instead.
- distinguish the sample size needed to ESTIMATE a quantity to a stated precision from the size needed to DETECT a difference of a stated size.
Assumes you have done Is this difference real? And is it big enough to care about? and Four suppliers, one answer, and the question it does not answer.
The problem
A packing team trialled a new changeover method on ten changeovers, against ten of the old one. The new method came out faster on average, the t-test returned a p-value above 0.05, and the trial was written up as "no significant improvement". The method was dropped and the project closed.
Their own numbers say the study had a 26% chance of detecting the difference they had come to find. The most likely outcome — by some margin — was exactly the result they got, whether the method worked or not. Nobody had asked how many changeovers it would take, and the calculation that would have answered it takes about ten seconds.
The idea
Every comparison has four quantities, and they are one equation:
- the difference worth detecting,
- the sample size,
- the significance level (alpha), the risk of calling a difference real when it is not,
- the power, the chance of finding the difference when it really is there.
Fix any three and the fourth is determined. That is the whole subject, and almost every mistake in it comes from fixing three by accident and never looking at the fourth.
Power is the one nobody fixes
Alpha gets chosen — nearly always 0.05, usually without discussion. The sample size gets chosen, by budget or by habit. The difference worth finding is usually implicit. Power is then whatever falls out, and nobody looks.
The convention is 80%: a one-in-five chance of missing a difference that is really there. That is not a law and it is not generous. It means a fifth of true improvements get written up as failures.
"Not significant" is not "no difference"
This is the sentence that costs the most.
A p-value above alpha means the data did not distinguish the difference from zero. It does not mean the difference is zero, and from an underpowered study it barely means anything at all — because a study with 26% power was always more likely to return that result than not, whether the effect was real or absent.
What can be said instead is the bound: the largest difference still consistent with the data. If that bound is smaller than anything worth acting on, "no difference" is a reasonable summary. If it is not, the honest summary is the study was too small to answer the question.
There is a worse thing than missing it
An underpowered study that DOES come out significant is not a small version of a good result. Only the larger sampling fluctuations clear the critical value, so the effects such a study detects are systematically overestimates.
The Hook's trial makes it concrete. At ten per group, with a standard deviation of 8 minutes, the smallest observed difference that could have come out significant is 7.5 minutes. The improvement they were looking for was 5. That study could not have reported the truth — it could only miss it, or overstate it by half again. Run enough small trials and the ones that get published are the ones that exaggerated.
The effect size is not a statistical quantity
The tool will ask what difference is worth detecting. Nothing in the data can supply it — it is a statement about what would change a decision. Five minutes off a changeover might be worth a project or worth nothing, depending on how many changeovers there are in a week and what else the line could do with the time.
Choosing it to make the sample size affordable inverts the whole exercise, and it is the most common way this calculation gets misused.
When the sample is already fixed — which is most of the time
Real projects rarely choose n. There were forty machines, or twelve weeks of history, or the batch is what it is. Asking "how many do I need" then produces a number nobody can act on.
Turn the equation round and ask what this data could possibly show. That question always has an answer, it does not depend on the result, and it usually ends the argument — because "this study can only see a difference twice the size of the one we care about" is a sentence about the team's own question rather than about statistics.
Worked example
The changeover trial from the Hook. Current changeovers average 46 minutes with a standard deviation of 8 minutes, and the team would act on an improvement of 5 minutes or more.
Step 1 — put the difference in standard deviations. The tests work in units of the process's own spread, so 5 minutes against a spread of 8 is
d = 5 / 8 = 0.625
That is the only translation step, and it is why the standard deviation has to be estimated first. A bad estimate propagates into the answer as its square.
Step 2 — the sample size. At 80% power and alpha = 0.05, two-sided:
42 changeovers per group — 84 in total.
The tool reports the power it achieves at 42 (80.8%) and at 41 (79.8%), so "the smallest n that works" is a checked claim rather than a rounding.
Step 3 — what the trial they ran could see. They ran ten per group. Against a 5-minute difference that is 26.3% power — the figure from the Hook. Turned round: the smallest difference ten per group could detect four times in five is
1.325 standard deviations = 10.6 minutes
They ran a study that could only see a 10.6-minute improvement, in order to find a 5-minute one. The result was not evidence about the method. It was evidence about the sample size.
Step 4 — read the curve, not the number. A single figure invites "can we do 40 instead". The curve answers it:
| Per group | Power |
|---|---|
| 10 | 26.3% |
| 20 | 48.7% |
| 30 | 66.3% |
| 40 | 78.8% |
| 42 | 80.8% |
| 60 | 92.4% |
| 90 | 98.6% |
| 120 | 99.8% |
Forty sits on the steep part, so the answer to "can we do 40" is nearly, and every extra changeover still buys something. Past 82 per group, ten more buy less than a single percentage point. That is where "collect more data" stops being an answer.
Step 5 — the question that is not this question. Estimating the current mean changeover time to within ±2.5 minutes takes 42 observations in total. Detecting a 5-minute difference takes 84 — 42 per group. Same process, same units, twice the data, and the second question only sounds like the first. Which of the two comes out larger depends entirely on how much precision you asked for: at ±2 minutes the estimate would need 64. Asking for "a big enough sample" without saying which question you mean is how a study ends up sized for neither.
Your turn
The tool opens on the changeover trial: a 5-minute difference, a standard deviation of 8, and the 10 per group they were about to run.
- Confirm the headline — 42 per group, 84 in total — and the panel saying that 10 per group can only detect 10.6 minutes with 26.3% power against the 5 you asked for.
- Raise the difference worth detecting to 10 minutes. The requirement falls to 12 per group, which the team could have afforded easily. This is the honest trade, and it is a decision about what matters rather than about statistics: are you willing to walk away from a 5-minute improvement in order to run a cheap study?
- Put it back to 5 and raise power to 95%. The requirement goes from 42 to 68 per group. Power is bought, not assumed.
- Set the standard deviation to 4 — the same process after its variation has been halved — and the requirement drops from 42 to 12 per group. Reducing variation makes every future study cheaper, which is an argument for control charts that has nothing to do with control charts.
- Now set the affordable sample to 200 and read the curve. The target line is crossed long before the right-hand edge, and the flattening note names the point past which more data buys almost nothing.
How many do you need?
You need 42 per group — 84 in total.
42 per group (84 in total) reaches 80.8% power against a difference of 0.625 standard deviations. 41 per group would give 79.8%.
Effect size · an input, not a result
0.625
Power at that n
80.8%
Power at 10 · what you were going to run
26.3%
What 10 per group can seeAt 10 per group this study would detect a difference of 10.6 minutes 80% of the time. You came here to detect 5 minutes. Against that difference it has 26.3% power — so a result of “no significant difference” would tell you nothing you did not already know.
The effect size of 0.625 is an input, not a result. Nothing in the data supplies it: it is a statement about how large a difference would have to be before anyone would act on it. Choosing it to make the sample size affordable inverts the whole exercise.
Estimating is not detecting. Knowing the mean to within ±2.5 minutes takes 42 observations in total; detecting a difference of 5 minutes takes 84 in total — 42 per group. They are different questions and they give different answers for the same study — asking for “a big enough sample” without saying which one you mean is how a trial ends up sized for neither.
A business decision, not a statistical one. Nothing in the data supplies it.
From history, a pilot, or a control chart. The answer scales with its square.
What the study can see at this size is the question worth asking when n is fixed.
How this is calculated
The sample size is solved from the noncentral t distribution — the distribution a t statistic actually follows when the null hypothesis is false — rather than from the normal approximation. The approximation gives 63 per group to detect half a standard deviation at 80% power where the exact answer is 64, and the reason for the gap is the point: it uses z where the test will use t, and t’s critical value depends on the n being solved for. The calculation is circular, so it is solved by search.
Every figure here is checked against published tables: 394, 64 and 26 per group for effect sizes of 0.2, 0.5 and 0.8, and 199, 34 and 15 for the one-sample design.
Source: Lenth, R.V. (1989), Algorithm AS 243, Applied Statistics 38(1); Cohen, J. (1988), Statistical Power Analysis for the Behavioral Sciences, 2nd ed.
Saved runs can be attached to a project deliverable as evidence. Both what you entered and what the tool computed are stored, so the result can be checked again later.
Check yourself
No hints. Wrong answers are explained, not softened.
A trial of 10 against 10 returns p = 0.31 for a difference the team would act on. The study had 26% power against that difference. What has been shown?
A team sized a study to detect a 5-minute difference where the standard deviation is 8 minutes, and got 42 per group. Somebody proposes detecting 2.5 minutes instead, keeping power and alpha the same. Roughly what happens to the sample size?
A project has 12 weeks of history and cannot get more. What is the useful question to ask of it?
After a non-significant result, an analyst computes power using the effect size the study observed, gets 26%, and reports it as evidence the study was too small. What is wrong with this?
A manager asks for 'a big enough sample to understand our changeover time'. What has to be settled before that can be answered?
A study is designed at alpha = 0.05 and 80% power to detect d = 0.5, and the sample size comes out at 64 per group. The sponsor will fund only 40 per group. Which of the other three quantities MUST change?
Worth remembering
What are the four quantities of a study design?
The difference worth detecting, the sample size, the significance level (alpha) and the power. They are one equation: fix any three and the fourth is determined. Power is the one nobody fixes, so it ends up wherever the other three left it.
Why is 'not significant' from a small study not evidence of no effect?
A p-value above alpha means the data did not distinguish the difference from zero. A study with 26% power was always more likely to return that result than not, whether the effect was real or absent — so it cannot tell those apart. Report the bound instead: the largest difference still consistent with the data.
How does sample size scale with the difference you want to detect?
With the INVERSE SQUARE. Halving the detectable difference roughly quadruples the sample — 42 per group becomes 162. Halving the standard deviation saves the same factor, which is why reducing variation makes every future study cheaper.
What is wrong with computing power from the effect you observed?
For a given test it is a strictly decreasing function of the p-value, so it carries nothing the p-value did not. Every non-significant result has low observed power, whether the true effect is enormous or zero. Ask instead what effect the design COULD have detected — that question has an answer and does not depend on the result.
Can you do this now?
Rate yourself honestly. We compare your rating with how you actually answered — the gap is more useful than either number alone.
I can state the four quantities a study design fixes — the difference worth detecting, the sample size, the significance level and the power — and solve for any one of them given the other three.
I can compute the sample size a comparison needs from the smallest difference worth acting on, and say why that difference is a business decision rather than a statistical one.
I can say what a study whose sample is already fixed by circumstance can and cannot detect, and decide on that basis whether it is worth running at all.
I can explain why a non-significant result from an underpowered study is uninformative rather than evidence that there is no effect.
I can show that power computed from the effect already observed adds nothing to the p-value, and name the question to ask instead.
I can distinguish the sample size needed to ESTIMATE a quantity to a stated precision from the size needed to DETECT a difference of a stated size.
Your rating is recorded alongside your drill results. Neither alone marks the competency as met.