Key points#
- The intended purpose defines the exact claim that the evidence must carry.
- EU IVDR clinical evidence combines three distinct elements: scientific validity, analytical performance, and clinical performance.
- A performance evaluation plan maps each claim to methods, data, acceptance logic, limitations, and remaining gaps.
- Evidence must match the users, population, specimens, setting, instrument, software, cutoff, and medical role described in the claim.
- The performance evaluation report and post-market performance follow-up keep the argument current throughout the device lifecycle.
Start by turning the intended purpose into testable statements#
An in vitro diagnostic, or IVD, can produce a technically precise result and still lack evidence for the medical use printed on its label, and the evaluation therefore begins with the intended purpose, not with a pile of study reports.
A sufficiently specific intended purpose identifies:
- the analyte, marker, or quantity detected or measured;
- the target condition, physiological state, or risk being assessed;
- whether the result is qualitative, semiquantitative, or quantitative;
- the specimen type and important collection conditions;
- the target population and relevant clinical spectrum;
- the intended user and use environment;
- the medical role, such as screening, aid to diagnosis, prognosis, prediction, treatment selection, or monitoring;
- important instrumentation, software, cutoffs, and interpretation rules.
Each detail can change the evidence question. Performance in serum does not automatically establish performance in capillary whole blood. A hospital-laboratory study may not support home use. Data from symptomatic referral patients may give a distorted account of screening performance in a low-prevalence population; an algorithm evaluated with one instrument version may not support a materially changed software pipeline.
The intended purpose is therefore the boundary condition for the evaluation. Broad wording increases the evidence burden rather than allowing evidence from a narrow setting to stretch farther.
Build a claim-to-evidence map before judging adequacy#
Article 56 and Annex XIII of the EU In Vitro Diagnostic Medical Devices Regulation organize performance evaluation around clinical evidence. MDCG 2022-2 describes three components that must work together:
- Scientific validity: the association of the analyte or marker with a clinical condition or physiological state.
- Analytical performance: the device's ability to detect or measure the target correctly under defined technical conditions.
- Clinical performance: the device's ability to produce results related to the target condition or state in the intended population and use.
These are connected, but they are not interchangeable. A well-established biomarker does not validate a particular assay. Low measurement imprecision does not establish diagnostic value. High sensitivity in a selected case-control sample does not necessarily predict performance in routine care.
Your claim-to-evidence map lists every material statement in the intended purpose and labeling. Beside each statement, it records the supporting source, method, and population. It records device version, result, uncertainty, applicability, and gap. This makes overreach visible. If no evidence row supports lay interpretation, for example, a claim of home use cannot be rescued by laboratory precision data.
Scientific validity asks whether the marker belongs in the medical question#
Scientific validity is the evidence that a marker is associated with the condition or physiological process named in the claim, and for an established analyte, authoritative consensus and a substantial literature may provide much of the support. For a novel marker, new data may be central.
The appraisal should specify a reproducible search, inclusion criteria, quality assessment, and relevance judgment. It should distinguish association from causation and describe conflicting results. Important applicability questions include:
- Is the marker defined in the same way as in the device claim?
- Does the association hold in the intended population and relevant subgroups?
- Does timing relative to symptoms, treatment, or disease stage change the marker?
- Are common comorbidities, medicines, or physiological states important modifiers?
- Does the proposed cutoff have scientific support, or was it selected after inspecting the same data used to report performance?
A marker can be scientifically valid for one medical role but not another: association with established disease, for example, does not by itself establish value for presymptomatic screening or treatment selection.
Analytical performance characterizes the complete measurement pathway#
Analytical studies test the device as a system: specimen, collection material, and transport. The system also covers preparation, reagents, and calibrators. It covers instrument, software, user actions, and result calculation. Relevant characteristics depend on the technology and claim. They can include trueness or bias, precision, and analytical sensitivity and specificity. They can include detection and quantitation limits, measuring interval, and linearity. They can include interference, cross-reactivity, and carryover. They can include hook effects, specimen stability, calibration, and metrological traceability.
Your design should challenge conditions that could plausibly affect use. Precision work may span days, lots, operators, instruments, and sites. Interference panels should reflect endogenous and exogenous substances relevant to the population. Specimen-stability claims need the stated temperatures, containers, freeze-thaw cycles, and transport intervals. For qualitative tests, performance near the cutoff deserves special attention because small analytical shifts can change the reported category.
Software is part of the measurement pathway when it converts signals, selects regions, applies a model, or combines markers. A fixed algorithm needs version control and verification. A model change can alter clinical performance even if the wet-laboratory assay is unchanged.
Analytical success should be interpreted against the medical claim. A bias that looks small in aggregate may be unacceptable near a decision boundary; conversely, a very low detection limit may have little medical value if concentrations below a higher threshold have no supported interpretation.
Clinical performance tests the device in its claimed context#
Clinical performance describes how results relate to the target condition or state in the intended population. Depending on the claim, measures can include diagnostic sensitivity and specificity, positive and negative predictive values, and likelihood ratios. They can include agreement, expected values, hazard discrimination, calibration, or other prespecified metrics.
No single metric travels unchanged across settings. Predictive values depend on prevalence. Spectrum differences can shift sensitivity and specificity when a study compares obvious cases with unusually healthy controls. Verification bias arises when only some participants receive the reference procedure. Imperfect reference standards can misclassify the very condition used to judge the device.
A credible clinical performance study therefore addresses:
- participant sampling and whether it represents the intended population;
- inclusion, exclusion, and recruitment timing;
- specimen collection and handling in the proposed workflow;
- blinding between index-test and reference results;
- the clinical reference standard and how disagreements are resolved;
- prespecified cutoffs and analysis populations;
- invalid, indeterminate, missing, and repeated results;
- sample-size assumptions and confidence intervals;
- important subgroups, sites, users, and instruments;
- any difference between the tested device and the marketed version.
Clinical performance is not the same as clinical utility. Showing that a test classifies a condition accurately does not necessarily show that using it improves patient outcomes or decisions, and if the intended purpose or risk-benefit argument depends on a downstream benefit, your evidence plan must address that additional question.
Decide whether existing evidence closes the gaps#
Your performance evaluation plan should inventory published literature, prior studies, routine-use data, comparable-device information where appropriate, and device-specific investigations. Each source needs appraisal for methodological quality, relevance, and contribution to the exact claim.
Existing evidence can be sufficient. A new clinical study is not valuable merely because it is new; the question is whether the available body of evidence is scientifically valid, complete, and applicable to the device and purpose. A study is needed when a material gap remains, such as a novel marker, unsupported population, new specimen, new user group, changed algorithm, uncertain cutoff, or missing comparison with an appropriate reference.
ISO 20916 describes good study-practice principles for IVD clinical performance studies using specimens from human participants. It addresses planning, responsibilities, and conduct. It addresses records, reporting, data credibility, and protection of participants and specimen donors. It does not replace the need to choose a design that answers the claim-specific question, and it does not cover analytical performance studies.
Write the report as a reasoned conclusion, not an archive#
The performance evaluation report should integrate the three evidence components. It explains the methods used to find and appraise evidence, presents results with uncertainty, resolves or preserves discordance, and states whether the body of evidence is sufficient for each claim. It should align with risk management, labeling, instructions for use, and the rest of the technical documentation.
A useful conclusion states boundaries. It may support one specimen but not another, one population but not children, trained laboratory users but not self-testing, or a narrow diagnostic-aid role rather than stand-alone diagnosis. Those limits are not editorial inconveniences. They are the honest perimeter of the evidence.
Traceability should also connect unresolved uncertainty to a control: further study, a warning, restricted use, monitoring, or a narrower claim. Simply labeling uncertainty as a limitation without a proportionate response leaves the evaluation incomplete.
Keep the evaluation synchronized with the marketed device#
Performance evaluation continues after market entry. Post-market performance follow-up can confirm assumptions in broader use, study underrepresented groups, monitor drift, investigate residual uncertainty, or evaluate a changed scientific context. Inputs can include follow-up studies, literature, and complaints. They can include vigilance reports, external quality assessment, and proficiency testing. They can include user feedback and real-world performance data.
Signals need a predefined route back into risk management and labeling. A changed disease definition, new treatment pathway, or altered prevalence can change the meaning of earlier evidence. So can a new interferent, specimen-handling trend, instrument update, or software modification. The appropriate result may be confirmation, a corrective action, or a new study. It may be a revised cutoff, added warning, or narrower intended purpose.
An IVD performance evaluation is strongest when every claim leads to evidence, every data set leads to an applicability judgment, and every important uncertainty leads to action, and that structure makes the clinical evidence reviewable at authorization and maintainable after launch.
Sources and further reading
- Regulation (EU) 2017/746 on In Vitro Diagnostic Medical Devices, especially Article 56 and Annex XIII, consolidated text (accessed 2026-07-15)
- European Commission MDCG 2022-2, Guidance on General Principles of Clinical Evidence for IVDs, January 2022 (accessed 2026-07-15)
- ISO 20916:2019, In Vitro Diagnostic Medical Devices, Clinical Performance Studies Using Specimens from Human Subjects (accessed 2026-07-15)
- European Commission, MDCG-Endorsed Guidance for In Vitro Diagnostic Medical Devices (accessed 2026-07-15)
Questions and answers
Is analytical accuracy enough to support an IVD clinical claim?
No. The evaluation must also support the marker-condition relationship and the device's performance for the intended medical purpose and population.
Does every IVD need a new prospective clinical performance study?
Not automatically. Existing evidence may be sufficient when it is reliable, applicable, and complete; unresolved claim-specific gaps determine whether new studies are needed.
When is a performance evaluation finished?
It is not a one-time document. Post-market evidence, scientific change, complaints, vigilance, and device modifications can require the evaluation and its conclusions to be updated.