Key points#
- Begin with one decision question that a multidisciplinary team can answer or reject.
- Define the context of use, including the model, data, scope, output, and other evidence in the decision.
- Rate model influence and the consequence of a wrong decision separately, then combine them as model risk.
- Distinguish model risk from model impact, which addresses departure from regulatory standards or expectations.
- Prespecify technical criteria that are proportional to the model's assigned role and stakes.
- Preserve the path from source data and code through evaluation, interpretation, and the final bounded claim.
MIDD is a chain of evidence, not a model category#
Model-informed drug development uses computational modeling and simulation to integrate data and knowledge for drug-development decisions. The methods can include population pharmacokinetics, physiologically based pharmacokinetics, concentration-response models, disease-progression models, model-based meta-analysis, quantitative systems pharmacology, agent-based models, and machine-learning approaches.
ICH M15 does not give any method a permanent evidentiary rank. A model may be reliable for one purpose and unsuitable for another. What matters is the full chain:
question -> context of use -> decision stakes -> evaluation -> model outcome -> multidisciplinary conclusion
Breaking any link weakens the claim. A good fit to historical data does not show that a model can predict a new population. A well-verified implementation does not show that its assumptions are biologically adequate. A plausible simulation does not establish that it should drive a high-stakes dose decision.
Write the question so it can fail#
The question of interest is the information MIDD is intended to provide for a development or regulatory decision; it should be explicit and specific enough for the evidence to be judged.
“Characterize pharmacokinetics” is an activity. “Determine whether the proposed regimen keeps systemic drug levels within a justified range for adults with reduced kidney function” is a decision question. It identifies a population, regimen, target, and conclusion that can be supported, qualified, or rejected.
Separate questions often require separate assessment tables. A model that supports trial sampling may not automatically support labeling. A model that predicts average concentrations may not support a claim about rare toxicity. Combining distinct questions behind a broad objective can hide differences in data, stakes, and evaluation needs.
An effective question states:
- the decision and development stage;
- the medicine, regimen, and population;
- the quantity or outcome to be estimated;
- the comparison or acceptable range;
- the action that follows each possible result.
Define exactly what the model is being asked to do#
The context of use describes the model's assigned job. Under M15, it should identify the model or linked models, their role and scope, the data used to build them, and the additional evidence that contributes to the answer.
This is more specific than naming a method. “Population PK model” does not tell you whether the model will describe variability, estimate a covariate effect, simulate a regimen, bridge between formulations, or support dosing in an understudied group.
A complete context of use should make the boundaries visible:
- input data and their provenance;
- represented and unrepresented populations;
- dose and concentration ranges;
- assumptions about adherence, disease, and co-medication;
- model outputs and uncertainty;
- other clinical or nonclinical evidence;
- intended contribution to the final decision;
- conditions under which the model should not be used.
The model is not “validated” for every future task. Evaluation is tied to this context.
Separate the four judgments that are often collapsed#
Model influence#
Model influence is the intended weight of model outcomes relative to other available information; it rises when the model is the sole or dominant basis for the decision and falls when strong independent evidence carries much of the conclusion. Influence concerns planned use, not whether the model later produces a convenient result. The rating and justification should be set early enough to shape the evaluation.
Consequence of a wrong decision#
This judgment considers the severity and likelihood of negative effects if the overall decision is wrong. Patient safety and lack of efficacy are central, but the analysis should be specific to the question and available information. An exploratory trial-design choice and a final regimen for a vulnerable population can have different consequences even when they use the same model.
Model risk#
M15 combines model influence with the consequence of a wrong decision to determine model risk. Low influence does not always make risk low if consequences are severe. High influence may be acceptable, but it usually requires stronger technical criteria and a more persuasive evaluation.
Model risk is not an intrinsic defect in modeling. It describes how much the model contributes to a possible wrong decision and its undesirable consequences for the particular question.
Model impact#
Model impact asks how far the proposed MIDD strategy differs from a regulatory standard, or from regulatory expectations where no formal standard exists. It supports planning and early discussion. It is distinct from the clinical stakes captured in model risk. Keeping these judgments separate improves communication. A familiar analysis can still support a consequential decision, and a scientifically novel approach can address a low-consequence internal choice.
Turn risk into prespecified technical criteria#
Technical criteria define what the model and its outputs must demonstrate to support the question. Choose them before you interpret the decisive results, and keep them proportional to model risk.
Criteria may address:
- data quality, completeness, and relevance;
- structural and statistical assumptions;
- parameter precision and identifiability;
- predictive checks within the relevant range;
- performance on external or withheld data;
- sensitivity to influential assumptions;
- stability under alternative models;
- uncertainty in simulated quantities;
- evidence near proposed decision boundaries.
Acceptance should not depend on a single familiar plot. Different errors can hide behind good average performance. A model intended to support a subgroup claim needs enough relevant data and diagnostics for that subgroup, and a model used to extrapolate beyond observed conditions needs a scientific argument for why its structure can carry that extrapolation.
Evaluate implementation, performance, and applicability#
M15 groups evaluation around verification, validation, and applicability assessment.
Verification#
Verification asks whether data-processing code, equations, software implementation, and calculations are correct. Useful controls include code review, independent reproduction of critical calculations, version locking, unit tests, data lineage, and checks against known results. Passing verification shows that the intended model was implemented. It does not show that the intended model is adequate.
Validation#
Validation examines performance and robustness against relevant observations. The methods depend on the modeling approach and can include goodness-of-fit diagnostics, predictive checks, residual analyses, external data, simulation-based checks, calibration, discrimination, and sensitivity analyses.
The evaluation should align with the decision. Accurate average concentration may be insufficient when the conclusion depends on the upper tail. Good predictions in typical adults may say little about children or people with organ impairment.
Applicability#
Applicability asks whether the data, assumptions, model, and evaluation are relevant to the context of use. It is possible for a technically sound model to fail this test because the target population, regimen, disease state, or concentration range lies outside the evidence, and applicability is where scientific transportability becomes explicit. The report should show which parts are interpolation supported by observed data and which are extrapolation resting on mechanistic assumptions or prior information.
A worked reasoning example#
Suppose your program asks whether a lower regimen is needed for adults with reduced kidney function.
The context of use assigns a population pharmacokinetic model to estimate how clearance and area under the concentration-time curve vary with a prespecified kidney-function measure, and a dedicated study contributes dense samples in a smaller group, while later trials add sparse samples across a wider population. Concentration-safety and concentration-efficacy information define a justified target range.
The decision worksheet should then ask:
- How much weight will the model carry compared with direct data?
- What harm could follow from selecting the wrong regimen?
- Are patients near the proposed kidney-function boundary represented?
- Are the function equation, body-size convention, and time window consistent across datasets?
- How sensitive is the conclusion to binding, active metabolites, and alternative covariate forms?
- Does uncertainty leave more than one regimen plausible?
The final claim should be narrow: the evaluated model and supporting evidence do or do not support a specified regimen for the represented population. It should not become a general statement that the model is correct.
Plan and report at two levels#
ICH M15 distinguishes the model analysis plan from the model analysis report. The plan records objectives, data, methods, assumptions, technical criteria, evaluation, and planned analyses. The report documents implementation, deviations, results, diagnostics, uncertainty, limitations, and conclusions.
Regulatory documentation then connects those technical records to the MIDD assessment. It should show how the key elements were rated, whether technical criteria were met, and whether the model outcomes qualify as MIDD evidence for the question.
This layered reporting prevents two common failures. The first is a polished regulatory summary that cannot be traced to code and data, and the second is a detailed technical report that never states what decision the work supports.
Review the claim, not just the model#
Before you accept a model-based conclusion, ask:
- Is the question explicit and decision-linked?
- Is the context of use bounded and complete?
- Are model influence, wrong-decision consequence, model risk, and model impact separately justified?
- Were technical criteria specified before final interpretation?
- Are verification, validation, and applicability all addressed?
- Does uncertainty reach the decision rather than ending at a confidence interval?
- Are unrepresented groups and extrapolations visible?
- Can the conclusion be reproduced from the analysis records?
- Was regulatory alignment sought early enough for the model's role and impact?
The strongest MIDD conclusion is precise. It says what a specific model, evaluated against specific criteria and considered with specific other evidence, can contribute to the one decision you have to make.
Sources and further reading
- FDA and ICH, M15 General Principles for Model-Informed Drug Development, Final Guidance, June 2026 (accessed 2026-07-15)
- FDA, M15 General Principles for Model-Informed Drug Development, Final Guidance PDF, June 2026 (accessed 2026-07-15)
- FDA, Population Pharmacokinetics, Final Guidance, February 2022 (accessed 2026-07-15)
- FDA, Newly Added Guidance Documents, M15 final status and release date (accessed 2026-07-15)
Questions and answers
Does ICH M15 rank one modeling method above all others?
No. Suitability depends on the question, context, data, model role, accepted scientific practice, and the evaluation needed for that use.
Is model risk the chance that software code contains an error?
No. In M15, model risk combines the model's influence on a decision with the possible consequence of a wrong decision; code verification is one part of model evaluation.
Can a model replace a clinical study?
Sometimes model-based evidence may change or reduce additional evidence needs, but only for a defined question with adequate evaluation and regulatory alignment.