A spirometry result is not interpreted in isolation. The measured volume and flow are compared with a reference distribution for people of similar age, sex, and height. For decades, many laboratories also selected different reference equations according to a patient's reported race or ethnicity. The American Thoracic Society now recommends a race-neutral average reference equation, commonly called GLI Global, instead of race-specific equations.
The important distinction is simple: you, your lungs, and the measured values do not change when the software changes. The predicted value, lower limit of normal, z-score, and sometimes the resulting category can change. That can affect clinical conversations and rules that rely on numerical cutoffs, so a transition needs more than a silent software update.
Key points#
- Race is a social and political classification, not a precise biological measurement of an individual's lungs.
- A reference equation estimates a distribution; it does not diagnose disease by itself.
- GLI Global removes the need to assign a patient to a race category for interpretation.
- Switching equations can reclassify results, especially near a lower limit or severity cutoff.
- Symptoms, test quality, trends, imaging, laboratory findings, and other clinical evidence remain essential.
What spirometry actually measures#
Spirometry commonly reports forced vital capacity, or FVC, and the volume exhaled in the first second, or FEV1. The FEV1/FVC ratio helps identify an obstructive pattern. These are measurements made during a coached maneuver. Acceptability and repeatability matter because an incomplete inhalation, early stop, cough, leak, or variable effort can distort them.
Interpretation adds a statistical comparison. A reference equation estimates the expected value and the spread around it for a person with specified characteristics, and a z-score describes how far the measurement lies from the reference mean in standard-deviation units. A commonly used lower limit of normal is the fifth percentile, near a z-score of -1.645.
This approach is more informative than applying a fixed rule such as 80% predicted to everyone. Normal variability changes with age and body size, so one percent-predicted cutoff can classify people unevenly. Even a statistically appropriate lower limit is not a biological cliff. Values just above and below it are similar measurements with different labels.
Why race appeared in older equations#
Reference equations were fitted to observed population data. Researchers found average differences among groups labeled by race or ethnicity and carried those labels into prediction; that may sound descriptive, but it treats broad, shifting social categories as if they were stable physiological variables.
Group averages can reflect many influences: childhood conditions, nutrition, air quality, housing, work-related hazards, smoking patterns, access to care, geography, ancestry, measurement methods, and selection of supposedly healthy reference samples. A race coefficient cannot identify which influence matters for one person. It can also normalize lower measured function in a population facing preventable harms, making the harm look like an expected baseline.
Race categories vary across countries, eras, institutions, and self-identification. People with mixed backgrounds do not fit a simple selector. The same individual could receive a different predicted value if a form or laboratory uses a different category. These are poor properties for a variable presented as individualized precision.
What the ATS recommendation changed#
The 2023 ATS official statement concluded that laboratories should move from race-specific interpretation toward a race-neutral average reference equation, and the society described GLI Global as the current practical option and emphasized continued study of patient outcomes and downstream consequences.
The recommendation is not a claim that one equation captures every healthy population perfectly. It is a judgment that requiring a race label is scientifically and ethically weaker than using a common reference while investigating more direct determinants of lung health; it also places the burden on any continued use of race to show benefit.
Race-neutral interpretation does not mean ignoring racism, unequal environmental conditions, or social determinants. Those factors belong in research, prevention, history-taking, and policy. Removing race from a prediction field should make clinicians more attentive to actual hazards and circumstances, not less.
How a report can change#
Suppose a person's FEV1 is 2.0 liters on two identical analyses. Under one reference equation, the predicted value might be lower and the z-score closer to the mean. Under another, the predicted value might be higher and the z-score farther below the mean. The measured 2.0 liters has not changed, but the statistical comparison has.
Population studies show that moving to race-neutral equations redistributes classifications. On average, some Black patients are classified as having more impairment than under older race-specific equations, while some White patients may be classified as having less. Effects differ by outcome, age, sex, and location within the reference distribution. Aggregate findings cannot tell you how your own report will change.
The COPDGene analysis evaluated how alternative approaches related to outcomes such as survival and respiratory events, and it supports the value of race-neutral z-score interpretation, but it remains an observational analysis in selected cohorts. No single cohort resolves every occupational, disability, transplant, treatment, or insurance threshold.
Diagnosis is larger than a reference equation#
An abnormal FEV1/FVC ratio can support an obstructive pattern, but a clinician still considers symptoms, bronchodilator response, smoking and work history, prior infections, medications, imaging, and alternative explanations. A low FVC on spirometry alone does not prove restriction; lung-volume testing may be needed. A normal value does not exclude asthma or another condition when symptoms are intermittent.
Likewise, severity labels do not directly equal symptom burden or prognosis. Two people with similar z-scores may differ in exercise capacity, oxygenation, exacerbations, imaging, and comorbid illness. A numerical category is one input to judgment, not a complete portrait.
Longitudinal comparison needs special care. If a laboratory changes equations between visits, the new z-score may differ even when the raw FEV1 and FVC are stable, and a report should spell out the equation and the software version used. Clinicians can compare raw values, review earlier data under the new equation when feasible, and distinguish a true physiological change from a reference change.
Thresholds have consequences#
Pulmonary function numbers can inform eligibility for treatment, surgery, rehabilitation, transplant evaluation, employment, compensation, disability programs, or research. Some policies use fixed percent-predicted categories that were created under older equations. Changing the equation without reviewing the policy can create unintended effects.
A responsible transition therefore includes validation of the laboratory system, staff education, clear report language, notice to referring clinicians, and an audit of high-stakes thresholds. Institutions should specify whether historical results will be recalculated and how apparent category changes will be communicated. Nobody should be told their lungs suddenly worsened when only the reference changed.
Policy makers should ask whether a threshold predicts an outcome that matters and whether it produces equitable decisions under the new reference, and where evidence is weak, a range plus clinical review may be safer than a rigid cutoff. The equation change is an opportunity to inspect the whole decision rule.
How to read a result during the transition#
Look for the equation name, predicted values, lower limits, and z-scores. Confirm that the test met quality criteria. Compare the raw FEV1, FVC, and ratio with prior measurements, not only the percent predicted. Ask whether your previous report used the same equation.
If a category changed, useful questions include:
- Did the measured value change, or only the reference?
- Is the difference large enough to exceed expected test variability?
- Does it fit symptoms and other evidence?
- Does a decision depend on a cutoff created for another equation?
- Would repeating a high-quality test or obtaining full pulmonary function testing change management?
These questions turn a surprising label into a traceable interpretation.
Limits of the evidence#
Reference populations are samples, not definitions of health. GLI Global pools information across groups and can still fit some geographic or demographic populations better than others. Studies of reclassification show what happens in their datasets; they do not establish every benefit or harm that will follow in every health system.
Outcome associations can support one interpretive approach without proving that all downstream decisions improve. Clinical thresholds may lag behind laboratory standards. Continued monitoring should examine accuracy, access, treatment, occupational effects, and patient understanding across populations.
The sound conclusion is neither that old reports were meaningless nor that a new equation is perfect. It is that race should not be used as a biological shortcut, statistical uncertainty should be visible, and decisions should rest on you and the full clinical record.
Sources and further reading
- American Thoracic Society official statement on race and pulmonary function test interpretation
- American Thoracic Society summary of its race-neutral recommendation
- COPDGene analysis of race-neutral spirometry interpretation
- ERS Breathe review of z-scores and percent predicted
- JAMA Network Open analysis of changes under race-neutral equations
Questions and answers
Did my lung function change when the laboratory adopted GLI Global?
Not because of the equation alone. The raw breathing measurement is the same; its predicted value and z-score may change. Compare raw values and test quality across dates.
Does race-neutral mean ancestry and environment never affect lung health?
No. Genetic variation, air quality, work hazards, housing, nutrition, infections, smoking, and unequal treatment can matter. A broad race label is not a precise measurement of those influences.
Is a value below the lower limit of normal a diagnosis?
No. It is a statistical flag. Diagnosis depends on the pattern, symptoms, history, test validity, and other clinical findings.
What should I ask if a threshold affects work, benefits, or treatment?
Ask which equation was used, whether the rule was validated with it, how close the value is to the cutoff, and whether clinical review or repeat testing is available. A clinician or pulmonary function laboratory can explain how a specific report was produced and how it should affect care.