“Accurate and precise” is common shorthand for a good measurement, but measurement science gives the words distinct jobs. Precision describes how closely repeated results agree under stated conditions. Trueness describes how close the average of many results is to a reference value. Accuracy describes closeness between a measured value and the value of the quantity being measured.
The vocabulary matters because a device can return nearly identical readings every time while carrying a systematic bias, and another method can average near a reference while producing readings too variable for a single-person decision. In clinical research, both problems can alter classification, treatment thresholds, effect estimates, and apparent reproducibility.
The vocabulary, carefully stated#
Measurement precision is closeness of agreement among indications or measured values obtained by replicate measurements on the same or similar objects under specified conditions; a standard deviation, coefficient of variation, or related statistic can quantify imprecision, but the conditions must be named.
Measurement trueness is closeness of agreement between the average of an effectively infinite number of replicate measured values and a reference quantity value. In practice, laboratories estimate this property using a suitable reference, enough repetitions, and a stated procedure. Bias is an estimate of systematic measurement error.
Measurement accuracy refers to closeness of agreement between a measured quantity value and a true quantity value of the measurand, though the formal vocabulary notes that accuracy is not itself a quantity and should not be reported as though it were a single numerical property. Statements such as “95 percent accurate” are ambiguous unless the author tells you how the number was calculated. Keeping the three words apart is what avoids the common mistake of using accuracy as a synonym for either low bias or high precision, and a sound report separates them.
Scenario 1: the tightly grouped but biased analyzer#
Imagine a laboratory analyzer tested repeatedly with a reference material assigned a glucose value of 100 units. The instrument returns 108, 108, 109, 108, and 109. The values are tightly grouped, so repeatability is strong. Their average remains meaningfully above the reference, so estimated bias is positive and trueness is poorer.
That pattern can arise from calibration error, reagent lot effects, or matrix differences. It can arise from environmental conditions or a method-specific interference. Repeating the test on the same system may create confidence because the same answer keeps appearing. Repetition does not remove a systematic error.
The corrective question is not “can it repeat?” but “what reference and traceability chain support the scale?” Calibration establishes the relationship between instrument indications and known standards under specified conditions. Metrological traceability connects a result through a documented, unbroken chain of calibrations, each contributing uncertainty, to an appropriate reference. Traceability does not mean the result is perfect. It makes the reference path explicit and allows uncertainty to be evaluated.
Scenario 2: the scattered method with a reassuring average#
Now imagine results of 88, 111, 97, 106, and 98 around the same reference. Their average may lie close to 100, suggesting reasonable estimated trueness across repetitions. The scatter is large, so any one reading may be poorly suited to a threshold decision.
This is why group-level bias alone does not establish individual-level performance. If a clinical cutoff separates categories, random variation can move a person back and forth even when the method is unbiased on average. Researchers should report within-run and between-run variation, not only mean difference. Averaging repeated readings can reduce random variation under suitable conditions, but it does not correct systematic bias and may not reflect ordinary use, so the evaluation should match the intended procedure: one result, duplicate measurements, or a longer series.
Repeatability and reproducibility depend on conditions#
Repeatability is precision under a set of conditions that usually keeps the procedure, operator, and measurement system the same. Those conditions also keep location and a short time interval the same. Reproducibility changes relevant conditions, which may include operators, laboratories, systems, and time.
Neither term means exact duplication. They describe variation under a declared condition set: a method can have excellent repeatability within one laboratory and poorer reproducibility across sites because training, calibration, environment, specimen handling, or equipment differs. Intermediate precision covers the conditions between those extremes: different days, different operators, or different instruments inside one laboratory. So what a report has to tell you is what changed, not just which label it used.
Scenario 3: a home device and a clinic method disagree#
A home measurement can differ from a clinic result even when both devices function as designed. The systems may use different specimen types, timing, algorithms, calibration hierarchies, or environmental assumptions. Transport and storage can affect one sample. The person's physiological state can change between readings.
Agreement studies should therefore examine paired measurements across the intended range and population. Correlation is not enough. Two methods can correlate strongly while one remains consistently higher. Difference plots, estimates of bias, limits of agreement, and clinically relevant error zones can show whether the disagreement matters.
Method comparison also needs a defensible reference procedure. Calling one routine device the “gold standard” does not make its value true. Its own uncertainty and limitations belong in the analysis.
Measurement error is not the same as biological variation#
A person's measured concentration can change because the biological quantity changed, because the collection conditions changed, or because the measurement procedure varied. Meals, posture, and time of day may contribute before the specimen reaches an analyzer. So may hydration, exercise, acute illness, and medication timing.
Analytical variation describes the measurement process. Preanalytical variation covers collection, handling, transport, and storage. Within-person biological variation is the real fluctuation around one person's homeostatic state, and between-person variation is the difference among people.
These components answer different questions. A perfectly repeatable analyzer cannot make a fluctuating biological marker stable. Conversely, invoking biology should not excuse a poorly controlled analytical system. Study protocols should standardize what can reasonably be standardized and report the remaining sources.
Scenario 4: a research outcome assembled from several measurements#
Clinical studies often turn a measurement into an endpoint. Blood pressure may be the mean of several readings, imaging may require a reader to trace a boundary, a questionnaire score may sum multiple items, and a model may transform raw sensor signals into an event count.
Each step can add variation or bias. Reader training can improve repeatability without ensuring that readers identify the correct structure. An automated segmentation method can be consistent but systematically miss a region. A questionnaire can yield stable scores while failing to measure the intended construct.
This introduces validity beyond metrology. Reliability asks whether the measure is consistent. Construct and criterion validity ask whether it represents what the study claims. Responsiveness asks whether it can detect meaningful change. A reproducible measurement of the wrong construct is not a useful endpoint.
Uncertainty is part of the result#
Measurement uncertainty is a nonnegative parameter characterizing the dispersion of quantity values attributed to a measurand based on the information used, and it integrates relevant sources rather than pretending the reported number is exact.
An uncertainty statement needs a model, inputs, assumptions, and coverage convention. It may include calibration uncertainty, repeatability, environmental corrections, reference values, and other contributions; the uncertainty attached to a laboratory or engineering measurement differs from a confidence interval around a study mean, although both express limited knowledge.
Uncertainty becomes especially important near a decision threshold. A reported value just above a cutoff does not transform uncertainty into certainty. The clinical interpretation may require confirmation, context, or a decision rule designed for borderline values.
Precision can increase apparent significance without fixing validity#
Large sample sizes can estimate the mean of a biased measurement very precisely. The confidence interval becomes narrow around the wrong value. Likewise, repeated readings can reduce standard error while preserving an invalid construct or systematic calibration difference.
This is one reason statistical precision and measurement precision should be kept distinct. Statistical precision concerns uncertainty in an estimated parameter. Measurement precision concerns agreement among repeated measured values under stated conditions. Better measurement can improve statistical analysis, but a narrow study confidence interval is not proof of an accurate instrument.
What to look for in a validation study#
A useful evaluation tells you:
- the measurand, including specimen, quantity, and conditions;
- intended users, setting, population, and measurement range;
- the reference material, method, or procedure and why it is suitable;
- calibration and traceability information;
- repeatability, intermediate precision, and reproducibility conditions;
- estimated bias across the range, not only at one point;
- uncertainty and handling of values near clinical cutoffs;
- preanalytical and biological sources of variation;
- missing, failed, censored, and out-of-range results;
- performance in relevant subgroups and ordinary workflows.
Acceptance criteria should come from the intended use. The allowable error for population surveillance may not be acceptable for titrating a narrow therapeutic index drug. A method is not simply valid or invalid in the abstract; it is fit or unfit for a defined purpose with stated consequences.
Sources and further reading
- Joint Committee for Guides in Metrology, International Vocabulary of Metrology, Measurement Accuracy
- Joint Committee for Guides in Metrology, Measurement Trueness
- Joint Committee for Guides in Metrology, Measurement Precision
- BIPM, Guides in Metrology
- NIST Technical Note 1297, Guidelines for Evaluating and Expressing Measurement Uncertainty
Questions and answers
Can a measurement be precise but inaccurate?
Yes. Repeated values can agree tightly while all are shifted by systematic bias. Calibration and comparison with an appropriate reference help detect the problem.
Does a result close to a reference prove the method is precise?
No. One result can land near a reference by chance. Precision requires replicate measurements under specified conditions.
Is reproducibility always more important than repeatability?
They answer different questions. Repeatability helps characterize the method under stable conditions. Reproducibility shows what happens when relevant conditions vary. Intended use determines which conditions matter. Clear measurement language prevents false reassurance. Precision, trueness, accuracy, uncertainty, and validity are not competing labels; together they show you how much confidence a result deserves, and for which decision.