Skip to content
Leantensify Learn

Green Belt · 15 min

Change everything at once, carefully

After this you can

  • set up a two-level factorial design and say what each run is and why every combination is needed.
  • explain why changing one factor at a time misses interactions, using a concrete pair of settings.
  • read an interaction and explain why the main effects involved cannot be recommended separately.
  • decide which effects in an unreplicated experiment stand out from noise, and say what kind of judgement that is.

Assumes you have done Four suppliers, one answer, and the question it does not answer.

The problem

An engineer improves a filtration process one factor at a time, holding everything else still. Starting from the baseline of 45 gal/h: raise the temperature — 71, a clear win, keep it high. Now try raising the formaldehyde concentration — 60, worse than 71. Conclusion: concentration hurts. Keep it low, move on.

Here is the same factor tried at low temperature: 45 without concentration, 68 with it. Concentration helps — by 23 gal/h.

Both experiments were run correctly. Concentration helps at low temperature and hurts at high temperature, and an experimenter who moves one factor at a time can only ever see the answer for the corner they happen to be standing in.

The idea

The one-factor-at-a-time instinct feels careful: change one thing, so you know what caused what. It has two defects, and neither is a matter of taste.

It cannot see interactions

An interaction is a factor whose effect depends on another factor's setting. They are not exotic — chemistry, machines and people are full of them: the glue that needs both heat and pressure, the setting that helps on one material and ruins another. When an interaction exists there is no such thing as "the effect of A". There is the effect of A at one setting of C, and a different one at the other, and any experiment that never varies them together cannot even represent the question.

It spends runs badly

Changing one factor at a time, each run informs one comparison. A 2ᵏ factorial — every combination of k factors at two levels each — makes every run inform every estimate: each effect is the difference between the average of half the runs and the average of the other half, a different halving each time. Sixteen runs of a 2⁴ estimate fifteen effects, each one backed by all sixteen observations. This is why the factorial is not a luxury for large budgets; it is the cheap option per question answered.

Reading the output

A main effect is the average change in the response between a factor's low and high settings. An interaction effect measures how much a factor's effect shifts with another factor. And the moment an interaction is real, the main effects inside it become averages of two different answers — the honest reading is the simple effects: the effect of A with C low, and with C high, quoted as a pair.

That has a practical consequence for recommendations: when A and C interact, the recommendation must name a pair of settings, never two separate ones.

The unreplicated problem

Most real factorials are run once per combination — runs are expensive. That has a price nobody advertises: with every degree of freedom spent on effects, there is no estimate of experimental error left. Nothing to compute an F test against.

The honest way out is Lenth's method: most effects in a sparse system are estimating nothing but noise, so the smaller effects themselves can serve as the noise floor. Effects that stand clear of that floor are worth believing; the rest are noise-sized. It is a screening judgement — good at pointing at what deserves a confirmation run, and not an F test. A tool that prints p-values for an unreplicated design is inventing them.

Two disciplines that cost nothing

Randomise the run order. The table is written in standard order; the process is run in random order, so a drift in the afternoon does not masquerade as the effect of whichever factor happened to change after lunch.

Confirm before acting. The winning combination is a prediction from a model. Run it again. If the confirmation lands where the model says, act; if not, the model just told you something the sixteen runs could not.

Worked example

A chemical product's filtration rate, with four candidate factors: temperature (A), pressure (B), formaldehyde concentration (C), stirring rate (D). A 2⁴ factorial — sixteen runs, every combination, one run each.

Step 1 — the effects. Each is a contrast of all sixteen runs:

EffectEstimateEffectEstimate
A (temp)21.625AC−18.125
B (pressure)3.125AD16.625
C (conc)9.875BC2.375
D (stirring)14.625AB0.125

Step 2 — which of these are real? One run per combination, so there is no error estimate and no F test. Lenth's method: the pseudo standard error comes out at 2.625, and the margin of error at 6.748. Five effects stand clear: A, C, D, AC and AD. Pressure — both its main effect and everything it touches — is noise-sized.

Step 3 — the interactions are the finding.

The AC interaction (−18.125) decomposed into simple effects:

Effect of temperature
at low concentration+39.75
at high concentration+3.5

Temperature transforms the process at low concentration and barely matters at high. The Hook is this row read from the other side: concentration helps at low temperature and hurts at high, and neither factor has a single "effect" to quote.

The AD interaction (+16.625) is the opposite shape: temperature is worth +5 at low stirring and +38.25 at high. So the answer to "should we raise the temperature?" is genuinely it depends — and the factorial can say on what.

Step 4 — the recommendation is a combination. High temperature, high stirring, low concentration, pressure wherever it is cheapest to run. Then a confirmation run at that combination, because the estimate is a prediction and sixteen runs bought a model, not a guarantee.

What one-factor-at-a-time would have reported from the same process: "temperature helps, concentration hurts" — the second half of which is true only at the corner it was measured from, and the AD interaction would never have been seen at all.

Dataset: ds-filtration-doe — the same data loads in the tool below, so you can reproduce every figure here yourself.

Your turn

The tool opens on the filtration experiment, with the design table in standard order beside the responses.

  1. Confirm the effects: Temp = 21.625, TempConc = −18.125, TempStir = 16.625, and Lenth's PSE of 2.625 with its margin of 6.748. The tool names effects by joining your factor names, so where the worked example above writes A, AC and AD — the textbook's letters — the screen reads Temp, TempConc and TempStir. Same contrasts, same numbers; rename a factor and every effect containing it renames with it.
  2. Read the half-normal plot. Ten effects lie on the dashed noise line; five peel away above it. That picture is the entire analysis of an unreplicated factorial.
  3. Read the simple-effects tables. Find the +39.75 / +3.5 split for temperature across concentration — the Hook's numbers, from the other side.
  4. Now load the plasma etch preset. It is replicated, so the same tool switches to real F tests and reports MS error on 8 degrees of freedom — the difference between screening and testing, visible in one click.
  5. In the etch data, find the gap effect: +52 at low power, −255 at high power. It reverses direction. Ask what "the main effect of gap is −101.6" could possibly mean to an operator, and notice the tool refuses to summarise it that way.
  6. Delete the last response line and read the error. A factorial with a missing run loses the orthogonality that makes its estimates independent — the tool refuses rather than quietly estimating something else.

Two-level factorial experiment

practice

5 effects distinguishable from noise — screening judgement, no error estimate exists

5 of 15 effects are distinguishable from noise (Temp, Conc, TempConc, Stir, TempStir), judged against Lenth's margin of 6.748 — a screening criterion, since an unreplicated design has no error estimate. The interactions TempConc, TempStir mean the factors involved cannot be set independently: the right setting for one depends on the other. The simple-effects table below shows each combination, and the recommendation has to be a PAIR of settings, not two separate ones. This is also exactly what one-factor-at-a-time experimentation cannot see. Before acting on this: run the winning combination again as a confirmation. An effect estimate is a prediction, and the confirmation run is where it meets the process.

Design
2^4, 1× replicated
Method
Lenth's method
Lenth PSE
2.625
Margin of error
6.748
Every effect this design can estimate, in standard order. An effect is the average change in the response between the low and high settings of its factors.
EffectEstimatevs margin 6.75Verdict
Temp21.625|21.625|Distinguishable
Pressure3.125|3.125|Noise-sized
TempPressure0.125|0.125|Noise-sized
Conc9.875|9.875|Distinguishable
TempConc-18.125|18.125|Distinguishable
PressureConc2.375|2.375|Noise-sized
TempPressureConc1.875|1.875|Noise-sized
Stir14.625|14.625|Distinguishable
TempStir16.625|16.625|Distinguishable
PressureStir-0.375|0.375|Noise-sized
TempPressureStir4.125|4.125|Noise-sized
ConcStir-1.125|1.125|Noise-sized
TempConcStir-1.625|1.625|Noise-sized
PressureConcStir-2.625|2.625|Noise-sized
TempPressureConcStir1.375|1.375|Noise-sized
Half-normal plot of 15 effect estimates. 5 effects peel away above the noise line: Conc at 9.88, Stir at 14.63, TempStir at 16.63, TempConc at 18.13, Temp at 21.63. The line's slope is the effect standard error, 2.625.noise lineConcStirTempStirTempConcTemp23.40Half-normal quantile|Effect|
Effects that estimate nothing fall on the dashed noise line; real ones peel away above it. Squares with labels are the effects the analysis calls distinguishable — the shape carries the meaning, not just the colour.

The interactions, decomposed — why one setting at a time is not a recommendation

Effect of Temp is 39.750 with Conc low, and 3.500 with Conc high.

Mean response at each combination of Temp and Conc.
Conc lowConc high
Temp low45.2573.25
Temp high85.0076.75

Effect of Temp is 5.000 with Stir low, and 38.250 with Stir high.

Mean response at each combination of Temp and Stir.
Stir lowStir high
Temp low60.2558.25
Temp high65.2596.50

This design is unreplicated, so there is no estimate of experimental error — every degree of freedom went into effects. Significance here comes from Lenth's method, which builds a noise floor from the smaller contrasts themselves. It is a screening judgement, good at finding the effects worth a confirmation run; it is not an F test, and a borderline effect deserves the replicated follow-up rather than a verdict.

Temp and Conc interact: the effect of Temp is 39.750 with Conc low and 3.500 with Conc high. The main effects of Temp and Conc average those two answers and describe neither; any recommendation has to name both settings together.

Temp and Stir interact: the effect of Temp is 5.000 with Stir low and 38.250 with Stir high. The main effects of Temp and Stir average those two answers and describe neither; any recommendation has to name both settings together.

Factor names
Run these in RANDOM order in real life; enter them here in this order.
#RunTempPressureConcStir
1(1)
2a+
3b+
4ab++
5c+
6ac++
7bc++
8abc+++
9d+
10ad++
11bd++
12abd+++
13cd++
14acd+++
15bcd+++
16abcd++++

Replicates on the same line, separated by commas.

Load a published experiment

How this is calculated
  • Effects are contrasts: the average response with the effect’s factors high minus the average with them low — effect = contrast / (n·2k−1), SS = contrast² / (n·2k). The contrasts are orthogonal, which is what lets 2k runs estimate 2k−1 effects independently.
  • Replicated designs test each effect with F on 1 and 2k(n−1) degrees of freedom against the pooled error.
  • Unreplicated designs have no error estimate — every degree of freedom went into effects. Significance uses Lenth’s pseudo standard error: PSE = 1.5 × median of the |effects| below 2.5 × (1.5 × median|effect|), with margin t(1−α/2, m/3) × PSE. It is a screening criterion, and the tool labels it as one rather than presenting it as an F test.
  • The half-normal plot ranks |effects| against half-normal quantiles. Effects that are pure noise fall on a line through the origin; real ones peel away above it.
  • Simple effects decompose each significant interaction, because the main effects it involves are averages of two different answers and cannot be quoted alone.

Source: Montgomery, D.C., Design and Analysis of Experiments 8e, Ch. 6 (Examples 6.1 and 6.2 are the presets, and the verification fixtures); Lenth, R.V. (1989), Technometrics 31(4), 469–473; Daniel, C. (1959) on the half-normal plot.

Check yourself

No hints. Wrong answers are explained, not softened.

A team optimises oven temperature holding line speed still, then optimises line speed at the new temperature, and locks in both. What has this procedure assumed?

A factorial reports: effect of additive = +12 at low temperature, −9 at high temperature. The manager asks for 'the effect of the additive' for the report. What is the honest answer?

An unreplicated 2⁴ finds one effect of 9.2 against Lenth's margin of 6.7. What kind of claim is 'this effect is significant'?

A 2³ factorial needs runs at 8 combinations, but run 'bc' was skipped to save time. What is lost?

Worth remembering

What can a factorial see that one-factor-at-a-time cannot?

Interactions — factors whose effect depends on another factor's setting. OFAT only ever measures a factor at the settings where the others are parked, so a conditional effect looks like a fixed one measured at whichever corner you stood in.

What happens to a main effect when its interaction is significant?

It becomes the average of two different answers — the effect at each setting of the interacting factor — and describes neither. Report the simple effects as a pair, and make any recommendation name both settings together.

Why does an unreplicated factorial need Lenth's method?

Every degree of freedom went into effects, so no error estimate exists and no F test is possible. Lenth builds a noise floor from the smaller effects themselves — a screening judgement that shortlists effects for a confirmation run, not a significance test.

Why randomise the run order of a designed experiment?

The design table is in standard order, and anything that drifts over time — warm-up, tool wear, the afternoon shift — would otherwise be credited to whichever factor happens to change in step with it. Random order turns time into noise instead of a phantom effect.

Can you do this now?

Rate yourself honestly. We compare your rating with how you actually answered — the gap is more useful than either number alone.

  • I can set up a two-level factorial design and say what each run is and why every combination is needed.

  • I can explain why changing one factor at a time misses interactions, using a concrete pair of settings.

  • I can read an interaction and explain why the main effects involved cannot be recommended separately.

  • I can decide which effects in an unreplicated experiment stand out from noise, and say what kind of judgement that is.

Your rating is recorded alongside your drill results. Neither alone marks the competency as met.