Green Belt · 15 min
Change everything at once, carefully
After this you can
- set up a two-level factorial design and say what each run is and why every combination is needed.
- explain why changing one factor at a time misses interactions, using a concrete pair of settings.
- read an interaction and explain why the main effects involved cannot be recommended separately.
- decide which effects in an unreplicated experiment stand out from noise, and say what kind of judgement that is.
Assumes you have done Four suppliers, one answer, and the question it does not answer.
The problem
An engineer improves a filtration process one factor at a time, holding everything else still. Starting from the baseline of 45 gal/h: raise the temperature — 71, a clear win, keep it high. Now try raising the formaldehyde concentration — 60, worse than 71. Conclusion: concentration hurts. Keep it low, move on.
Here is the same factor tried at low temperature: 45 without concentration, 68 with it. Concentration helps — by 23 gal/h.
Both experiments were run correctly. Concentration helps at low temperature and hurts at high temperature, and an experimenter who moves one factor at a time can only ever see the answer for the corner they happen to be standing in.
The idea
The one-factor-at-a-time instinct feels careful: change one thing, so you know what caused what. It has two defects, and neither is a matter of taste.
It cannot see interactions
An interaction is a factor whose effect depends on another factor's setting. They are not exotic — chemistry, machines and people are full of them: the glue that needs both heat and pressure, the setting that helps on one material and ruins another. When an interaction exists there is no such thing as "the effect of A". There is the effect of A at one setting of C, and a different one at the other, and any experiment that never varies them together cannot even represent the question.
It spends runs badly
Changing one factor at a time, each run informs one comparison. A 2ᵏ factorial — every combination of k factors at two levels each — makes every run inform every estimate: each effect is the difference between the average of half the runs and the average of the other half, a different halving each time. Sixteen runs of a 2⁴ estimate fifteen effects, each one backed by all sixteen observations. This is why the factorial is not a luxury for large budgets; it is the cheap option per question answered.
Reading the output
A main effect is the average change in the response between a factor's low and high settings. An interaction effect measures how much a factor's effect shifts with another factor. And the moment an interaction is real, the main effects inside it become averages of two different answers — the honest reading is the simple effects: the effect of A with C low, and with C high, quoted as a pair.
That has a practical consequence for recommendations: when A and C interact, the recommendation must name a pair of settings, never two separate ones.
The unreplicated problem
Most real factorials are run once per combination — runs are expensive. That has a price nobody advertises: with every degree of freedom spent on effects, there is no estimate of experimental error left. Nothing to compute an F test against.
The honest way out is Lenth's method: most effects in a sparse system are estimating nothing but noise, so the smaller effects themselves can serve as the noise floor. Effects that stand clear of that floor are worth believing; the rest are noise-sized. It is a screening judgement — good at pointing at what deserves a confirmation run, and not an F test. A tool that prints p-values for an unreplicated design is inventing them.
Two disciplines that cost nothing
Randomise the run order. The table is written in standard order; the process is run in random order, so a drift in the afternoon does not masquerade as the effect of whichever factor happened to change after lunch.
Confirm before acting. The winning combination is a prediction from a model. Run it again. If the confirmation lands where the model says, act; if not, the model just told you something the sixteen runs could not.
Worked example
A chemical product's filtration rate, with four candidate factors: temperature (A), pressure (B), formaldehyde concentration (C), stirring rate (D). A 2⁴ factorial — sixteen runs, every combination, one run each.
Step 1 — the effects. Each is a contrast of all sixteen runs:
| Effect | Estimate | Effect | Estimate | |
|---|---|---|---|---|
| A (temp) | 21.625 | AC | −18.125 | |
| B (pressure) | 3.125 | AD | 16.625 | |
| C (conc) | 9.875 | BC | 2.375 | |
| D (stirring) | 14.625 | AB | 0.125 |
Step 2 — which of these are real? One run per combination, so there is no error estimate and no F test. Lenth's method: the pseudo standard error comes out at 2.625, and the margin of error at 6.748. Five effects stand clear: A, C, D, AC and AD. Pressure — both its main effect and everything it touches — is noise-sized.
Step 3 — the interactions are the finding.
The AC interaction (−18.125) decomposed into simple effects:
| Effect of temperature | |
|---|---|
| at low concentration | +39.75 |
| at high concentration | +3.5 |
Temperature transforms the process at low concentration and barely matters at high. The Hook is this row read from the other side: concentration helps at low temperature and hurts at high, and neither factor has a single "effect" to quote.
The AD interaction (+16.625) is the opposite shape: temperature is worth +5 at low stirring and +38.25 at high. So the answer to "should we raise the temperature?" is genuinely it depends — and the factorial can say on what.
Step 4 — the recommendation is a combination. High temperature, high stirring, low concentration, pressure wherever it is cheapest to run. Then a confirmation run at that combination, because the estimate is a prediction and sixteen runs bought a model, not a guarantee.
What one-factor-at-a-time would have reported from the same process: "temperature helps, concentration hurts" — the second half of which is true only at the corner it was measured from, and the AD interaction would never have been seen at all.
Dataset: ds-filtration-doe — the same data loads in the tool below, so you can reproduce every figure here yourself.
Your turn
The tool opens on the filtration experiment, with the design table in standard order beside the responses.
- Confirm the effects: Temp = 21.625, TempConc = −18.125, TempStir = 16.625, and Lenth's PSE of 2.625 with its margin of 6.748. The tool names effects by joining your factor names, so where the worked example above writes A, AC and AD — the textbook's letters — the screen reads Temp, TempConc and TempStir. Same contrasts, same numbers; rename a factor and every effect containing it renames with it.
- Read the half-normal plot. Ten effects lie on the dashed noise line; five peel away above it. That picture is the entire analysis of an unreplicated factorial.
- Read the simple-effects tables. Find the +39.75 / +3.5 split for temperature across concentration — the Hook's numbers, from the other side.
- Now load the plasma etch preset. It is replicated, so the same tool switches to real F tests and reports MS error on 8 degrees of freedom — the difference between screening and testing, visible in one click.
- In the etch data, find the gap effect: +52 at low power, −255 at high power. It reverses direction. Ask what "the main effect of gap is −101.6" could possibly mean to an operator, and notice the tool refuses to summarise it that way.
- Delete the last response line and read the error. A factorial with a missing run loses the orthogonality that makes its estimates independent — the tool refuses rather than quietly estimating something else.
Two-level factorial experiment
practice5 effects distinguishable from noise — screening judgement, no error estimate exists
5 of 15 effects are distinguishable from noise (Temp, Conc, TempConc, Stir, TempStir), judged against Lenth's margin of 6.748 — a screening criterion, since an unreplicated design has no error estimate. The interactions TempConc, TempStir mean the factors involved cannot be set independently: the right setting for one depends on the other. The simple-effects table below shows each combination, and the recommendation has to be a PAIR of settings, not two separate ones. This is also exactly what one-factor-at-a-time experimentation cannot see. Before acting on this: run the winning combination again as a confirmation. An effect estimate is a prediction, and the confirmation run is where it meets the process.
- Design
- 2^4, 1× replicated
- Method
- Lenth's method
- Lenth PSE
- 2.625
- Margin of error
- 6.748
| Effect | Estimate | vs margin 6.75 | Verdict |
|---|---|---|---|
| Temp | 21.625 | |21.625| | Distinguishable |
| Pressure | 3.125 | |3.125| | Noise-sized |
| TempPressure | 0.125 | |0.125| | Noise-sized |
| Conc | 9.875 | |9.875| | Distinguishable |
| TempConc | -18.125 | |18.125| | Distinguishable |
| PressureConc | 2.375 | |2.375| | Noise-sized |
| TempPressureConc | 1.875 | |1.875| | Noise-sized |
| Stir | 14.625 | |14.625| | Distinguishable |
| TempStir | 16.625 | |16.625| | Distinguishable |
| PressureStir | -0.375 | |0.375| | Noise-sized |
| TempPressureStir | 4.125 | |4.125| | Noise-sized |
| ConcStir | -1.125 | |1.125| | Noise-sized |
| TempConcStir | -1.625 | |1.625| | Noise-sized |
| PressureConcStir | -2.625 | |2.625| | Noise-sized |
| TempPressureConcStir | 1.375 | |1.375| | Noise-sized |
The interactions, decomposed — why one setting at a time is not a recommendation
Effect of Temp is 39.750 with Conc low, and 3.500 with Conc high.
| Conc low | Conc high | |
|---|---|---|
| Temp low | 45.25 | 73.25 |
| Temp high | 85.00 | 76.75 |
Effect of Temp is 5.000 with Stir low, and 38.250 with Stir high.
| Stir low | Stir high | |
|---|---|---|
| Temp low | 60.25 | 58.25 |
| Temp high | 65.25 | 96.50 |
This design is unreplicated, so there is no estimate of experimental error — every degree of freedom went into effects. Significance here comes from Lenth's method, which builds a noise floor from the smaller contrasts themselves. It is a screening judgement, good at finding the effects worth a confirmation run; it is not an F test, and a borderline effect deserves the replicated follow-up rather than a verdict.
Temp and Conc interact: the effect of Temp is 39.750 with Conc low and 3.500 with Conc high. The main effects of Temp and Conc average those two answers and describe neither; any recommendation has to name both settings together.
Temp and Stir interact: the effect of Temp is 5.000 with Stir low and 38.250 with Stir high. The main effects of Temp and Stir average those two answers and describe neither; any recommendation has to name both settings together.
| # | Run | Temp | Pressure | Conc | Stir |
|---|---|---|---|---|---|
| 1 | (1) | − | − | − | − |
| 2 | a | + | − | − | − |
| 3 | b | − | + | − | − |
| 4 | ab | + | + | − | − |
| 5 | c | − | − | + | − |
| 6 | ac | + | − | + | − |
| 7 | bc | − | + | + | − |
| 8 | abc | + | + | + | − |
| 9 | d | − | − | − | + |
| 10 | ad | + | − | − | + |
| 11 | bd | − | + | − | + |
| 12 | abd | + | + | − | + |
| 13 | cd | − | − | + | + |
| 14 | acd | + | − | + | + |
| 15 | bcd | − | + | + | + |
| 16 | abcd | + | + | + | + |
Replicates on the same line, separated by commas.
Load a published experiment
How this is calculated
- Effects are contrasts: the average response with the effect’s factors high minus the average with them low — effect = contrast / (n·2k−1), SS = contrast² / (n·2k). The contrasts are orthogonal, which is what lets 2k runs estimate 2k−1 effects independently.
- Replicated designs test each effect with F on 1 and 2k(n−1) degrees of freedom against the pooled error.
- Unreplicated designs have no error estimate — every degree of freedom went into effects. Significance uses Lenth’s pseudo standard error: PSE = 1.5 × median of the |effects| below 2.5 × (1.5 × median|effect|), with margin t(1−α/2, m/3) × PSE. It is a screening criterion, and the tool labels it as one rather than presenting it as an F test.
- The half-normal plot ranks |effects| against half-normal quantiles. Effects that are pure noise fall on a line through the origin; real ones peel away above it.
- Simple effects decompose each significant interaction, because the main effects it involves are averages of two different answers and cannot be quoted alone.
Source: Montgomery, D.C., Design and Analysis of Experiments 8e, Ch. 6 (Examples 6.1 and 6.2 are the presets, and the verification fixtures); Lenth, R.V. (1989), Technometrics 31(4), 469–473; Daniel, C. (1959) on the half-normal plot.
Check yourself
No hints. Wrong answers are explained, not softened.
A team optimises oven temperature holding line speed still, then optimises line speed at the new temperature, and locks in both. What has this procedure assumed?
A factorial reports: effect of additive = +12 at low temperature, −9 at high temperature. The manager asks for 'the effect of the additive' for the report. What is the honest answer?
An unreplicated 2⁴ finds one effect of 9.2 against Lenth's margin of 6.7. What kind of claim is 'this effect is significant'?
A 2³ factorial needs runs at 8 combinations, but run 'bc' was skipped to save time. What is lost?
Worth remembering
What can a factorial see that one-factor-at-a-time cannot?
Interactions — factors whose effect depends on another factor's setting. OFAT only ever measures a factor at the settings where the others are parked, so a conditional effect looks like a fixed one measured at whichever corner you stood in.
What happens to a main effect when its interaction is significant?
It becomes the average of two different answers — the effect at each setting of the interacting factor — and describes neither. Report the simple effects as a pair, and make any recommendation name both settings together.
Why does an unreplicated factorial need Lenth's method?
Every degree of freedom went into effects, so no error estimate exists and no F test is possible. Lenth builds a noise floor from the smaller effects themselves — a screening judgement that shortlists effects for a confirmation run, not a significance test.
Why randomise the run order of a designed experiment?
The design table is in standard order, and anything that drifts over time — warm-up, tool wear, the afternoon shift — would otherwise be credited to whichever factor happens to change in step with it. Random order turns time into noise instead of a phantom effect.
Can you do this now?
Rate yourself honestly. We compare your rating with how you actually answered — the gap is more useful than either number alone.
I can set up a two-level factorial design and say what each run is and why every combination is needed.
I can explain why changing one factor at a time misses interactions, using a concrete pair of settings.
I can read an interaction and explain why the main effects involved cannot be recommended separately.
I can decide which effects in an unreplicated experiment stand out from noise, and say what kind of judgement that is.
Your rating is recorded alongside your drill results. Neither alone marks the competency as met.