Skip to content
Leantensify Learn

How many do we need — and if we cannot get that many, what can this study actually see?

How many observations a comparison needs — and, when the sample is already fixed, what it could possibly show you. The question worth asking before the trial is booked rather than after it.

Sample Size and Power · measure · Master Black Belt

Use this when

  • Before booking a trial, so the number is argued about in advance rather than defended afterwards
  • When the sample is already fixed by circumstance — 40 machines, 12 weeks of history — and the useful question is what it could possibly detect
  • When a trial came back 'no significant difference' and somebody is about to read that as 'the change did not work'

How many do you need?

You need 42 per group — 84 in total.

42 per group (84 in total) reaches 80.8% power against a difference of 0.625 standard deviations. 41 per group would give 79.8%.

Effect size · an input, not a result

0.625

Power at that n

80.8%

Power at 10 · what you were going to run

26.3%

What 10 per group can seeAt 10 per group this study would detect a difference of 10.6 minutes 80% of the time. You came here to detect 5 minutes. Against that difference it has 26.3% power — so a result of “no significant difference” would tell you nothing you did not already know.

Power against sample size. The curve rises from 7% at 2 to 100% at 126. It reaches the 80% target at 42. Beyond about 82, ten more observations add less than one percentage point.0%25%50%75%100%2126observations per group
Target met at 42. Beyond 82, ten more observations buy under one point of power.

The effect size of 0.625 is an input, not a result. Nothing in the data supplies it: it is a statement about how large a difference would have to be before anyone would act on it. Choosing it to make the sample size affordable inverts the whole exercise.

Estimating is not detecting. Knowing the mean to within ±2.5 minutes takes 42 observations in total; detecting a difference of 5 minutes takes 84 in total — 42 per group. They are different questions and they give different answers for the same study — asking for “a big enough sample” without saying which one you mean is how a trial ends up sized for neither.

A business decision, not a statistical one. Nothing in the data supplies it.

From history, a pilot, or a control chart. The answer scales with its square.

What the study can see at this size is the question worth asking when n is fixed.

How this is calculated

The sample size is solved from the noncentral t distribution — the distribution a t statistic actually follows when the null hypothesis is false — rather than from the normal approximation. The approximation gives 63 per group to detect half a standard deviation at 80% power where the exact answer is 64, and the reason for the gap is the point: it uses z where the test will use t, and t’s critical value depends on the n being solved for. The calculation is circular, so it is solved by search.

Every figure here is checked against published tables: 394, 64 and 26 per group for effect sizes of 0.2, 0.5 and 0.8, and 199, 34 and 15 for the one-sample design.

Source: Lenth, R.V. (1989), Algorithm AS 243, Applied Statistics 38(1); Cohen, J. (1988), Statistical Power Analysis for the Behavioral Sciences, 2nd ed.

Saved runs can be attached to a project deliverable as evidence. Both what you entered and what the tool computed are stored, so the result can be checked again later.

How this is calculated

Sample size and power for the one- and two-sample t tests are solved from the NONCENTRAL t distribution, which describes the behaviour of a t statistic when the null hypothesis is false, rather than from the normal approximation. Lenth's Algorithm AS 243 computes the noncentral t as a Poisson mixture of incomplete beta functions with a running bound on the unsummed tail, and refuses rather than returning a partial sum if that bound is not met. The exact calculation is circular — it uses the t critical value, which depends on the n being solved for — so the size is found by bracketing and bisection, and the result reports the power at n-1 as well as at n, making minimality a checked claim rather than an assertion. The same relationship is solved in the other two directions: the smallest detectable effect at a fixed sample, by bisection on the effect size, and the sample needed to estimate a mean to a stated half-width, iterated on t rather than z because the interval the study will actually report uses t. Two-proportion sizing uses the normal approximation, which is what published tables for proportions use, and every result carries whether it is exact or approximate. Effect size is an input throughout and is labelled as one: it is a statement about what size of difference would justify acting, and nothing in the data supplies it.

Source: Lenth, R.V. (1989), Algorithm AS 243: Cumulative Distribution Function of the Non-Central t Distribution, Applied Statistics 38(1), 185-189; Cohen, J. (1988), Statistical Power Analysis for the Behavioral Sciences 2nd ed. — the effect-size definitions and the conventions, quoted as conventions rather than as thresholds; Hoenig, J.M. & Heisey, D.M. (2001), The Abuse of Power, The American Statistician 55(1), 19-24, on why power computed after the result answers nothing.

Learn the method