Evidence explainer

Evidence and research methods

How a Patient-Reported Outcome Measure Earns Trust: Reading It Through COSMIN

A questionnaire does not earn trust with a familiar name and a high Cronbach alpha. COSMIN starts somewhere harder: do these items represent this concept, for these people?

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Define what the score is supposed to represent
  2. Content validity comes first
  3. Structural validity asks how items form scores
  4. Internal consistency is not “the higher the better”
  5. Reliability and measurement error answer different questions
  6. Construct validity requires explicit hypotheses
  7. Cross-cultural validity and measurement invariance
  8. Criterion validity requires a real criterion
  9. Responsiveness is validity of a change score
  10. Interpretability and feasibility complete the decision
  11. Reading a COSMIN systematic review
  12. A practical appraisal sequence

A patient-reported outcome measure, or PROM, asks people directly about symptoms, function, well-being, or another health concept without interpretation by a clinician or observer. The direct voice is valuable, but a questionnaire can still ask the wrong questions, combine incompatible concepts, fluctuate too much, miss meaningful change, or perform differently across languages and groups. COSMIN provides a framework for evaluating those risks property by property.

Define what the score is supposed to represent#

The same instrument can be suitable for one purpose and poor for another, and a scale developed to compare average fatigue in rheumatoid arthritis trials may not be accurate enough to make decisions about one person with cancer. A short screening tool may not measure treatment benefit. A translated version may not retain the same meaning.

State the construct, target population, setting, language, administration method, and use. Is the aim to discriminate between people at one time, evaluate change, predict an event, or support an individual threshold? Which recall period applies? Who completes it, and are proxy responses allowed?

The score's measurement model matters. In a reflective model, items are manifestations of an underlying construct; changes in the construct are expected to influence item responses. Internal consistency and factor structure are relevant. In a formative index, distinct items jointly define a construct; they need not correlate, so Cronbach alpha may be inappropriate.

Content validity comes first#

COSMIN treats content validity as the most important measurement property. Three questions guide it:

Strong development combines literature and theory with qualitative work involving people from the target population and relevant professionals, and cognitive interviews can reveal that a term is interpreted differently, a response option cannot represent lived experience, or an item assumes an activity some participants never perform. Expert panels alone are insufficient for patient understanding. A large statistical validation cannot repair missing content: factor analysis only analyzes the items that developers chose to include.

Structural validity asks how items form scores#

Structural validity evaluates whether item responses reflect the dimensional structure the scoring assumes; if a total score combines physical function, emotional well-being, and pain, evidence should support one overall dimension or justify separate subscales.

Classical test theory commonly uses confirmatory factor analysis. Item response theory or Rasch models add assumptions about item behavior, local independence, monotonicity, and model fit. Exploratory analysis can generate a structure but should ideally be confirmed in new data.

Fit statistics should not be cherry-picked. Sample size, estimator for ordinal responses, missing-data handling, correlated residuals, and post hoc modifications affect conclusions. A model made to fit through many data-driven changes may not generalize.

Internal consistency is not “the higher the better”#

Internal consistency describes relationships among items in a scale. Cronbach alpha is often reported, but a high value can result from many redundant items. It does not prove unidimensionality, validity, stability, or responsiveness.

COSMIN interprets internal consistency after there is at least low-certainty evidence for sufficient structural validity in a reflective scale. Otherwise, alpha may summarize a mixture of concepts. Values that are extremely high can suggest repeated wording rather than broad content coverage. Report the statistic for each unidimensional subscale, not only the full questionnaire. Ordinal items may require methods suited to categorical data rather than assumptions designed for continuous values.

Reliability and measurement error answer different questions#

Reliability is the proportion of observed-score variance attributable to true differences among people rather than measurement error. Test-retest reliability often uses an intraclass correlation coefficient; the interval between administrations should be long enough to reduce recall but short enough that the construct is unlikely to change, and participants should be clinically stable.

Reliability depends on population variability. The same amount of measurement noise yields a lower intraclass correlation in a homogeneous group than in a diverse group. So the coefficient is not a universal property of the instrument.

Measurement error is expressed in score units, such as the standard error of measurement and smallest detectable change, and a smallest detectable change estimates how large a person's change must be to exceed random error with a specified level of confidence.

To judge sufficiency, compare error with a minimal important change or another clinically justified threshold, and if the smallest detectable change is larger than the smallest change people consider important, the instrument may distinguish groups while being unreliable for interpreting individual improvement.

Construct validity requires explicit hypotheses#

When no credible gold standard exists, construct validity is tested through hypotheses about relationships with other measures and differences between known groups. A pain-interference scale might correlate strongly with another interference measure, moderately with pain intensity, and weakly with an unrelated construct.

Those direction and magnitude hypotheses should be stated before anyone looks at the results. A paper that tells you only that “most correlations were significant” has shown very little, because significance tracks sample size rather than the predicted pattern. Known-groups validity also needs defensible groups that genuinely should differ. Using a clinical classification partly based on the PROM creates circular evidence.

Cross-cultural validity and measurement invariance#

Translation requires more than word substitution. Concepts, idioms, daily activities, and response styles differ. Good adaptation uses forward and backward translation or other structured methods, expert review, cognitive testing, and documentation.

Measurement invariance asks whether people with the same underlying construct respond similarly across groups such as language, age, sex, or culture. Differential item functioning can reveal an item that behaves differently after accounting for the construct. If invariance fails, group comparisons may reflect item behavior rather than health differences. Researchers might revise items, use group-specific parameters, remove problematic items, or avoid the comparison.

Criterion validity requires a real criterion#

Criterion validity compares a PROM with a defensible gold standard. For many patient-reported constructs, no gold standard exists; another questionnaire is not automatically one. Agreement with an imperfect instrument is better framed as construct validity.

A long-form instrument can sometimes serve as a reference for a short form derived from it, but shared items can inflate agreement. Diagnostic accuracy methods are appropriate only when the reference classification is valid for the intended use and threshold.

Responsiveness is validity of a change score#

Responsiveness is the ability to detect change in the construct over time. It is not shown by a significant within-group p-value. Large samples can make tiny changes significant, and uncontrolled improvement can reflect natural history, regression to the mean, co-interventions, or expectation.

Good studies prespecify hypotheses about change-score correlations or differences among groups expected to improve, remain stable, or worsen. The external anchor should itself measure change credibly and be interpretable. Effect size and standardized response mean describe magnitude relative to variability; they do not establish that change is important or caused by treatment. In a randomized trial, the between-group estimate remains central for treatment effect.

Interpretability and feasibility complete the decision#

Interpretability is not a measurement property in the COSMIN taxonomy, but it determines whether scores make sense. Look for score distributions, floor and ceiling effects, reference values, minimal important change, and thresholds with uncertainty.

A minimal important change is not one fixed universal constant. It can differ by baseline severity, direction of change, population, anchor, method, and individual versus group use. The smallest detectable change and minimal important change answer different questions.

Feasibility includes respondent burden, reading level, accessibility, licensing cost, administration time, scoring, missing-item rules, training, and compatibility with routine workflow. A psychometrically strong instrument that people cannot complete or clinicians cannot score may be the wrong choice.

Reading a COSMIN systematic review#

The 2024 COSMIN guideline version 2.0 sets out a structured review: define the question and eligibility, search, extract instrument characteristics, assess studies with the Risk of Bias checklist, rate each property against criteria, synthesize results, grade certainty, and formulate recommendations.

Do not average all measurement properties into one score. A study can use excellent methods and find that a property is insufficient. Conversely, favorable results from a very doubtful study are weak evidence. Method quality, result sufficiency, and certainty are separate judgments. Recommendations should match use. An instrument may be recommended, potentially useful pending further evidence, or not recommended because high-certainty evidence shows an important property is insufficient.

A practical appraisal sequence#

Name the PROM version, language, construct, population, setting, and intended decision. Read the development and content-validity work before you look at a single coefficient. Confirm the scoring structure, then interpret internal consistency. Compare reliability and measurement error with the intended individual or group use.

Check prespecified construct and responsiveness hypotheses, invariance across relevant groups, missing-data rules, floor and ceiling effects, and important-change estimates. Finally, consider burden, licensing, accessibility, and whether a better-validated alternative exists.

Sources and further reading

  1. COSMIN, Improving the Selection of Outcome Measurement Instruments
  2. Mokkink and colleagues, COSMIN Guideline for Systematic Reviews of PROMs Version 2.0, Quality of Life Research (2024)
  3. Prinsen and colleagues, COSMIN Guideline for Systematic Reviews of Patient-Reported Outcome Measures, Quality of Life Research (2018)
  4. Terwee and colleagues, COSMIN Methodology for Evaluating Content Validity of PROMs, Quality of Life Research (2018)
  5. Gagnier and colleagues, COSMIN Reporting Guideline for Studies on Measurement Properties of PROMs Version 2.0 (2024)

Questions and answers

Does a Cronbach alpha above 0.9 prove a good questionnaire?

No. It can reflect redundant items and says nothing by itself about content, dimensionality, stability, change detection, or clinical usefulness.

Is test-retest correlation enough for reliability?

Usually not. Ordinary correlation measures association, not agreement. An intraclass correlation chosen for the design, along with score-level measurement error, is generally more informative.

What is the difference between smallest detectable change and minimal important change?

Smallest detectable change concerns measurement noise. Minimal important change concerns a change considered meaningful. Both are needed to judge individual change.

Can a PROM be valid in one language but not another?

Yes. Translation, culture, administration, and item behavior can change meaning. Each version and intended group needs relevant evidence.

Is responsiveness the same as treatment efficacy?

No. Responsiveness asks whether the measure detects real change. Treatment efficacy requires a valid comparison, usually between randomized groups, and cannot be inferred from change in one group alone.