Skip to content
Leantensify Learn

Yellow Belt · 15 min

Measuring the gauge before you measure the process

After this you can

  • judge whether a measurement system is contributing more variation than the process it measures.

Assumes you have done Capable, or just in control? and Two people, one measurement, different answers.

The problem

A team spent four months reducing variation in a shaft diameter. The control chart wandered, the capability study came back poor, and every improvement they tried moved the numbers a little and never enough. Somebody eventually asked two operators to measure the same shaft. They disagreed by more than the entire tolerance band the team had been trying to hold. Thirty-eight per cent of the variation the team had been attacking was the hand-held gauge, and no amount of work on the process could have removed it.

The idea

Every number you have ever been shown came out of a measurement system, and that system has its own variation. When it is large relative to the part-to-part variation, your control chart, your capability study and your improvement project are all describing the gauge as much as the process — and nothing downstream can recover from it.

The observed variation is the sum of two variances:

σ²observed = σ²process + σ²measurement

Variances add, standard deviations do not. That is why the measurement system in the worked example below, contributing "only" 14.44% of the variance, shows up as 38.00% of the standard deviation — 10 x sqrt(14.44) — and why the two percentage columns on a Gage R&R report are constantly confused.

The two components of measurement variation

Repeatability — the same person, the same part, the same gauge, twice. Variation here is the equipment: resolution, play, how firmly it seats. Sometimes called equipment variation.

Reproducibility — different people, the same part, the same gauge. Variation here is the method: how the part is held, where it is measured, what "seated correctly" means to each person. Sometimes called appraiser variation, which invites blaming the appraiser; it is almost always an operational definition that was never written down.

Gage R&R is the two together, and the study that separates them tells you which problem you have. Large repeatability means buy a better gauge. Large reproducibility means write the method down — much cheaper, and far more common.

There is a third term worth knowing about: appraiser × part interaction, where different people measure different parts differently. That usually means the measurement is harder on some parts than others — a feature that is awkward to reach on the large sizes, for example.

Reading the report

Two percentage columns, and they answer different questions.

%Contribution is a variance ratio. The components sum to exactly 100%, so it is the one that can honestly be put in a pie chart.

%StudyVar is a standard deviation ratio and it does not sum to 100%, because the sum of square roots is not the square root of the sum. Two components at 50% contribution each read 70.71% here and total 141%. The exact relationship is:

%StudyVar = 10 × √(%Contribution)

Read one against the other's thresholds and you will reject a good gauge or accept a bad one. The conventional bands (under 10% acceptable, over 30% unacceptable) are stated on %StudyVar; the same gauge on %Contribution would be under 1% and over 9%.

ndc — number of distinct categories = 1.41 × (part-variation SD / Gage R&R SD), truncated. It is the number of non-overlapping groups of parts the system can actually tell apart. Below 5 the gauge can do little more than sort parts into "high" and "low" — it cannot support process analysis, whatever the percentages say.

And note the numerator: part variation, not total. Total always exceeds part variation, so substituting it always flatters the gauge.

The order of operations

A Gage R&R comes before the capability study and before the control chart, not after them, because both of those are built on numbers this gauge produced. A team that runs the study last — as in the Hook — spends the intervening months improving a process that was never the problem.

Worked example

Ten parts, three appraisers, three trials each: ninety measurements of shaft diameter in millimetres. That design is the standard one, and it is standard because smaller studies give estimates too wide to act on.

The interaction, first. F(18, 60) = 1.51, p = 0.119. AIAG's rule pools the interaction into error when p > 0.25; here it is below that, so the interaction is retained — appraisers do not measure all parts the same way. Reproducibility is therefore the appraiser term plus the appraiser × part term.

Which path was taken matters and is usually invisible on a report form: pooling changes the error estimate, and therefore every percentage below it.

The components:

Source% Contribution (variance, sums to 100)% Study var (std dev, does not)
Repeatability9.93%31.50%
Appraiser2.83%16.81%
Appraiser × part1.69%12.98%
Part-to-part85.56%92.50%
Gage R&R14.44%38.00%

Check the identity on the Gage R&R row: 10 × √14.44 = 10 × 3.80 = 38.00. Exactly.

The verdict: 38.00% study variation is unacceptable. The conventional band puts anything over 30% there.

ndc = 3 (3.43 before truncation). Below the floor of 5. This gauge can sort the ten parts into about three groups; it cannot tell you where inside those groups a part sits.

Where the problem is. Repeatability at 9.93% contribution against appraiser plus interaction at 4.52% — so the larger share is the equipment, not the people. That points at the gauge itself: resolution, wear, how the part seats. Had the split been the other way round, the answer would have been to write the method down, which costs an afternoon rather than a purchase order.

And what this means for the four months in the Hook. The team's capability study and control chart were both computed from numbers this gauge produced, so 14.44% of the variance they were attacking never lived in the process at all. Every improvement that genuinely worked was partly masked by a gauge that could not see it.

Dataset: ds-gage-rr-study — the same data loads in the tool below, so you can reproduce every figure here yourself.

Your turn

The tool opens with the full ninety-measurement study.

  1. Confirm Gage R&R at 38.00% study variation and 14.44% contribution, and check the identity yourself: 10 × √14.44.
  2. Note the partition column sums to exactly 100% and the study-variation column does not. That is not a rounding artefact — it is the difference between adding variances and adding standard deviations.
  3. Enter a tolerance range of 4 mm. %Tolerance appears and is far worse than %StudyVar, because the tolerance is narrower than the part spread in this study. Which metric you accept against is a real decision, and it should be made before you see the numbers.
  4. Delete appraisers B and C, leaving only A. The tool refuses, and the message is the lesson: with one appraiser there is no reproducibility component and no interaction, so a crossed Gage R&R is not defined at all — it names the simple repeatability study you would have to run instead. It will not report a GRR with a fabricated appraiser variation of zero, which is what a single-appraiser study quietly produces when a tool lets it.
  5. Delete the third trial from every row. Repeatability is now estimated from two readings instead of three, and the tool tells you the study is below the recommended design.

Gage R&R (crossed ANOVA)

practice

Not acceptable 38.00% %GRR of study variation (standard deviation ratio)

38.00% exceeds 30% — unacceptable. The measurement system needs improvement before the data it produces can support a decision.

AIAG MSA 4th ed., Ch. II §D

Gage R&R · % study var

38.00%

Part-to-part · % study var

92.50%

ndc · min 5

3

Variance components with their contribution to total variance and to total study variation.
SourceVariance% Contributionvariance · sums to 100% Study varstd dev · does NOT sum to 100
Repeatability (equipment variation, EV)0.244899.93%31.50%
Reproducibility: appraiser (AV)0.069742.83%16.81%
Reproducibility: appraiser x part interaction0.041581.69%12.98%
Part-to-part (PV)2.1110985.56%92.50%
Gage R&Rrepeatability + reproducibility0.3562114.44%38.00%

The two percentage columns are not the same quantity. %Contribution is a variance ratio and the partition sums to exactly 100%. %Study var is a standard deviation ratio and does not — the sum of square roots is not the square root of the sum. The identity between them is %StudyVar = 10 × √(%Contribution), so a component at 30% contribution reads 54.8% study variation. Comparing one against the other’s thresholds is the most common misreading of a Gage R&R report.

ndc = 3 (3.43 before truncation) — the number of distinct groups of parts this system can actually tell apart. Below 5, so the system can do little more than sort parts into "high" and "low" — it cannot support process analysis.

Interaction: Appraiser*part interaction F(18,60) = 1.5094, p = 0.1187 <= alpha = 0.25. The interaction was RETAINED: appraisers do not measure all parts the same way. Reproducibility is the appraiser term plus the appraiser*part term.

ndc = 3 (raw 3.43); AIAG MSA 4th ed. Ch. II §D requires at least 5. The system cannot reliably distinguish enough groups of parts to support process analysis.

No tolerance range supplied, so %Tolerance was not computed. When a measurement system is used for product acceptance, %Tolerance — not %StudyVar — is the metric AIAG applies the acceptance bands to.

part, appraiser, trial1, trial2, trial3 — one row per combination. Every appraiser measures every part the same number of times.

10 parts × 3 appraisers × 3 trials

USL − LSL. Supply it for %Tolerance; leave it blank and %Tolerance is simply not reported rather than guessed.

How this is calculated

A random-effects two-factor crossed ANOVA with replication. Expected mean squares give the variance components; the interaction is tested against error and pooled into it when p > 0.25, which is AIAG’s rule. Which path was taken is stated above rather than left implicit.

  • Repeatability (equipment) = MSerror.
  • Reproducibility (appraiser) = appraiser + appraiser×part components.
  • Gage R&R = repeatability + reproducibility.
  • Study variation multiplier: AIAG MSA 4th ed. (2010): 6 sigma, 99.73% of a normal distribution. Minitab default.. It cancels out of %StudyVar and %Contribution and ndc — both are ratios — and scales %Tolerance, the absolute EV/AV/GRR/PV/TV study-variation figures.
  • ndc = 1.41 × (part-variation SD / Gage R&R SD), truncated. The numerator is PART variation, not total: total always exceeds part variation, so substituting it always flatters the gauge.

Source: AIAG, Measurement Systems Analysis 4th ed., Ch. III §B (ANOVA method) and Ch. II §D (ndc); Montgomery, D.C., Introduction to Statistical Quality Control 7e, §8.7.

Saved runs can be attached to a project deliverable as evidence. Both what you entered and what the tool computed are stored, so the result can be checked again later.

Check yourself

No hints. Wrong answers are explained, not softened.

A Gage R&R reports 25% study variation and 6.25% contribution. A colleague says the two figures are inconsistent. Are they?

A study shows repeatability at 5% contribution and reproducibility at 22%. What does that point at, and what would you do?

A capability study on a new process returns Cpk = 0.9. The gauge has never had an MSA. What is the first thing to do?

Worth remembering

What is the difference between repeatability and reproducibility?

Repeatability is the same person measuring twice — the equipment. Reproducibility is different people measuring the same part — the method. Large reproducibility usually means the operational definition was never written down.

Why do %Contribution and %StudyVar disagree on a Gage R&R report?

One is a variance ratio and sums to 100%; the other is a standard deviation ratio and does not. %StudyVar = 10 × √(%Contribution), exactly. The acceptance bands differ between the two scales.

What does ndc below 5 mean?

The gauge can distinguish fewer than five groups of parts — enough to sort them high and low, not enough for process analysis. It uses PART variation over Gage R&R variation; substituting total variation always flatters the gauge.

Can you do this now?

Rate yourself honestly. We compare your rating with how you actually answered — the gap is more useful than either number alone.

  • I can judge whether a measurement system is contributing more variation than the process it measures.

Your rating is recorded alongside your drill results. Neither alone marks the competency as met.