Evidence explainer

Health policy, systems, and equity

SaMD: Analytical Versus Clinical Validation

Medical software needs a valid clinical association, correct and reliable processing, and performance that supports its intended use in the target population.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Start with the intended use
  2. Valid clinical association asks whether the output matters
  3. Analytical validation asks whether the software works as specified
  4. Clinical validation asks whether performance supports the purpose
  5. The reference standard can limit all three claims
  6. Subgroups and sites are part of the intended population
  7. Human factors can change clinical performance
  8. Validation continues after release
  9. A claim-to-evidence table
  10. References

“Validated” is incomplete when it appears without a claim, version, population, setting, and endpoint. For Software as a Medical Device, or SaMD, the International Medical Device Regulators Forum separates clinical evaluation into three connected questions. Is the software output linked to the clinical condition or physiological state it claims to address? Does the software correctly and reliably transform its inputs into that output? Does the output achieve its intended purpose in the target population and care context?

Those questions are called valid clinical association, analytical validation, and clinical validation. Passing one does not establish the others. A flawless calculation can implement an irrelevant biomarker. A clinically meaningful predictor can be coded incorrectly. Strong retrospective accuracy can fall apart when input quality, users, prevalence, or workflow changes.

Start with the intended use#

You cannot judge the evidence until the product claim is precise. An intended-use statement should identify the user, target population, setting, input, output, and decision the software supports. “Analyzes images” is a function. “Prioritizes adult chest radiographs with suspected pneumothorax for radiologist review in an emergency worklist” is closer to a testable clinical claim.

Each phrase creates an evidence obligation. “Adult” sets an age boundary. “Chest radiographs” defines image types and acquisition conditions. “Suspected pneumothorax” needs an operational reference standard. “Prioritizes” makes queue behavior and delay relevant. “Radiologist review” preserves a human role that must be represented in usability and workflow studies.

The same algorithm can support different products. An educational display, a triage notification, and an autonomous diagnostic output have different consequences when wrong. IMDRF's risk framework considers both the seriousness of the healthcare situation and how significant the software's information is to a decision; validation depth should follow the actual claim and risk, not the sophistication of the code.

Valid clinical association asks whether the output matters#

A valid clinical association is a supported relationship between the SaMD output and the targeted clinical condition or physiological state. Evidence may come from peer-reviewed literature, professional guidance, secondary analyses, or new clinical research, and the route depends on how established the relationship is and how closely existing evidence matches the proposed use.

Suppose you are reviewing software that calculates an already accepted clinical index from standard inputs. Published evidence and guidelines may support the association between that index and the condition; a new image-derived surrogate that claims to predict deterioration needs a more direct evidence program because both the measurement and its clinical meaning are new.

Association does not mean causation. A feature can correlate with an outcome without being a treatment target. Nor does an association justify every decision. A marker associated with future risk may support stratification but not prove that treating people above one threshold improves outcomes, and the claimed use sets how far the evidence must travel.

Evidence should define the endpoint, time horizon, comparator, population, and uncertainty. If the association comes from adults in specialty hospitals, it does not automatically support pediatric primary care. If the literature studied laboratory measurements, it may not support a software-derived proxy without bridging evidence.

Analytical validation asks whether the software works as specified#

IMDRF uses analytical validation for evidence that software correctly processes input data and generates accurate, reliable, and precise output. This is sometimes called analytical or technical performance. It includes more than checking that a formula returns the expected answer once.

For deterministic calculations, testing can cover formula accuracy, units, rounding, boundary conditions, invalid values, missing inputs, and reproducibility across supported systems. For image or signal algorithms, the program may include sensitivity, specificity, localization accuracy, repeatability, and performance across devices, acquisition protocols, compression levels, and artifacts. For machine-learning functions, dataset construction and independence are also central.

The test set must represent the intended input domain and remain separate from development and tuning. If you keep checking a held-out set while changing the model, you have turned it into another development set. Leakage can occur through duplicate patients, related images, preprocessing fitted on all data, or labels influenced by the output being evaluated.

Input quality deserves explicit acceptance criteria. A model can perform well on curated records and fail when fields are stale, units differ, images are cropped, or sensors drop data. Safe behavior may include rejecting unusable input, reporting the reason, requesting a new acquisition, or routing the task to a manual process. An unsupported confident output is not an acceptable fallback.

Performance summaries need denominators and intervals. An overall area under a curve can conceal poor sensitivity at the operating threshold, a mean error can conceal dangerous tails, and an aggregate accuracy figure can conceal a subgroup or a device type with sparse data. Test the metrics that connect to the claim and the harm analysis.

Clinical validation asks whether performance supports the purpose#

Clinical validation evaluates whether the output achieves its intended purpose in the target population and care context. Appropriate measures depend on the claim. A diagnostic function may require sensitivity, specificity, predictive values, and failure rates against a suitable reference standard. A risk model needs calibration, discrimination, and decision-relevant performance at the stated horizon. A triage tool may need time-to-review, missed urgent findings, queue displacement, and user response.

This stage is not simply analytical validation with patients added. It tests the clinically meaningful relation between output and intended use. Prevalence, disease spectrum, treatment patterns, user behavior, and workflow can change what a technically accurate result accomplishes.

Clinical validation also does not always prove clinical utility: a study can show that software detects a condition accurately without showing that its use improves outcomes, shortens delay, reduces unnecessary procedures, or causes acceptable harms. When the product claims improvement in care, or when user response is a major part of benefit and risk, a comparative prospective workflow study may be needed.

The strongest design depends on the question. A retrospective study may be appropriate for a locked image classifier's technical and clinical performance. A prospective silent study can test data flow and output stability without changing care, while a randomized or carefully controlled implementation study may be needed to estimate what happens when users act on the output.

The reference standard can limit all three claims#

Clinical labels are measurements, not pure truth. Pathology may be sampled, expert panels may disagree, a diagnosis code may reflect billing rather than a prespecified phenotype, and a future outcome can be missing altogether because the person received care somewhere else.

A validation report should identify how the reference was defined, who assessed it, whether assessors knew the software output, the timing between measurements, and how disagreements were resolved. When no single standard exists, the study may use adjudication, longitudinal follow-up, a composite, or latent-class methods. Each approach has assumptions.

Training against a flawed label can create a product that reproduces the label rather than the condition, so the high agreement in the report is real and also narrower than the clinical claim it is being used to support. Documentation should let you tell “predicts the recorded code” from “diagnoses the disease.” Those are different products.

Subgroups and sites are part of the intended population#

Overall performance can look acceptable while errors concentrate in people underrepresented in development data, and the relevant groups can involve age, sex, skin tone, language, comorbidity, disease severity, device model, care setting, and the site itself. Which of those matter should follow plausible mechanisms and intended use, not a fixed demographic checklist. The overall figure hides every one of them.

Subgroup estimates need enough cases and noncases for precision. A table of point estimates from tiny cells creates false reassurance. Prospective sample-size planning, confidence intervals, and a plan for unresolved uncertainty are more useful. Sometimes the correct result is that a claim must remain narrower until evidence grows.

Multisite evaluation tests more than geography. It adds variation in workflow, equipment, prevalence, documentation, and users. A random train-test split within one pooled dataset may allow site-specific patterns to appear on both sides. Holding out sites or time periods can provide a more credible transport test, although no one study proves universal performance.

Human factors can change clinical performance#

Users do not receive raw metrics. They see a screen, alert, explanation, threshold, and recommended next action. Poor interface design can turn a sound output into delayed care, automation bias, duplicated work, or alert fatigue.

Human-factors evaluation asks whether intended users understand the output and limitations, identify invalid inputs, respond correctly to critical alerts, and recover during downtime. It should include foreseeable stress, interruptions, and atypical cases. Training cannot substitute for a design that makes common dangerous errors easy.

Automation level matters. A low-specificity tool may be acceptable as a silent second reader if every output receives expert review, yet unsafe if it interrupts a queue hundreds of times per day. A high-performing model may still cause harm if it displaces more urgent work or encourages users to ignore contradictory clinical information.

Validation continues after release#

Clinical evaluation is a lifecycle process. Data distributions change, clinical practice changes, devices and interfaces update, users adapt, and rare failures accumulate. Postmarket monitoring should connect premarket claims and hazards to measurable signals.

Useful measures can include input rejection, subgroup performance, false negatives, alert response, overrides, time-to-action, complaints, adverse events, cybersecurity incidents, and version adoption. Action thresholds and owners should be defined before a metric moves. Monitoring without a response plan is observation, not control.

An update can invalidate earlier evidence. A change to model weights, threshold, input source, population, user, interface, or intended purpose may require new analytical or clinical evaluation. FDA's final guidance on predetermined change control plans provides a pathway for certain planned modifications to AI-enabled device software functions, but it does not remove the need to specify, assess, and control the changes. Version traceability should connect source code, model artifact, training data, configuration, labeling, risk controls, test results, and released product, because “the model was validated” tells you nothing if no one can identify which model the evidence covered.

A claim-to-evidence table#

A practical review can be organized into one row per claim:

  1. exact intended-use statement and product version;
  2. clinical association and supporting sources;
  3. analytical endpoint, dataset, comparator, and acceptance criterion;
  4. clinical endpoint, target population, setting, and uncertainty;
  5. user interaction and failure pathway;
  6. subgroup and site evidence;
  7. residual limitations in labeling;
  8. postmarket metric and action threshold.

Gaps become visible. A triage claim with only area-under-the-curve evidence has no queue outcome. A diagnostic claim with no invalid-input denominator omits operational failures. A broad population with one-site data lacks transport evidence. The table turns “validated” into claims you can check.

Related articles cover regulation as a design input and monitoring model drift. The site's research overview connects these methods to responsible clinical evidence.

References#

  1. IMDRF N41 clinical evaluation framework
  2. FDA SaMD clinical evaluation guidance
  3. IMDRF N12 risk categorization framework
  4. FDA Clinical Decision Support Software guidance
  5. FDA guidance on predetermined change control plans

Questions and answers

Is analytical validation the same as software verification?

They overlap but are not identical labels in every framework. Verification checks requirements; analytical validation addresses whether the SaMD correctly processes inputs into accurate, reliable, and precise outputs. A complete program links both.

Does high accuracy prove clinical validation?

Only if the metric, threshold, reference standard, population, and setting support the intended clinical purpose. A high aggregate metric on a mismatched dataset is insufficient.

Must every SaMD have a randomized trial?

No. Study design should fit the claim and risk. A randomized study may be needed when the claim concerns clinical impact or when user response is central, but other questions can be answered by analytical, retrospective, or prospective performance studies.

Can published literature establish the clinical association?

Yes, when the association is well established and matches the output, population, and use. A novel surrogate or expanded claim may require original evidence.

Is a product permanently validated after authorization?

No. Evidence applies to a defined version and use. Product changes and real-world drift require continuing assessment, and some changes need additional regulatory review.