ICH E8(R1) defines the quality of a clinical study as fitness for purpose: a high-quality trial protects participants and generates information reliable enough to answer its research question and support decisions. Quality is not synonymous with perfect data, maximum monitoring, or a large collection of standard operating procedures.
The guideline's central method is quality by design. During planning, a multidisciplinary team identifies the factors critical to participant rights, safety, well-being, and the reliability and meaning of results, and it then prevents or controls errors that matter, in proportion to their probability, detectability, and impact. Retrospective review, monitoring, and audits remain useful. But they cannot rescue a question, endpoint, population, or workflow that was poorly designed.
Begin with the decision, not the process map#
A trial exists to answer a question. The first quality task is to state that question and the decision it will inform. The population, treatment strategies, and comparator should align with it. So should the endpoint, follow-up, and estimand.
If the objective is vague, operational perfection cannot make the result meaningful. A trial can collect every scheduled field and still fail because the endpoint does not matter to patients, the control no longer represents relevant care, or follow-up is too short for the intended claim.
Fitness for purpose is contextual. An early pharmacology study and a confirmatory outcomes trial support different decisions and need different precision, duration, participants, and controls, so the same data issue can be critical in one study and negligible in another. That is why quality by design starts before the final protocol, while the research question, development stage, available knowledge, and intended use of the result can still shape the design.
What makes a factor critical to quality#
A critical-to-quality factor is an attribute whose integrity is fundamental to protecting participants or producing reliable and meaningful results; it should trace directly to the research question, important risks, or interpretability.
Common examples include:
- eligibility criteria that define the target disease and risk group;
- informed-consent processes and protection of vulnerable participants;
- randomization and allocation concealment;
- masking where knowledge could bias decisions or assessment;
- correct intervention assignment and accounting;
- timely recognition and reporting of important safety events;
- an endpoint definition that is clinically meaningful and measured consistently;
- retention and outcome ascertainment sufficient to prevent informative missingness;
- central laboratory or imaging procedures that determine the primary endpoint;
- data lineage and system controls for variables driving the primary analysis;
- prespecified analysis choices tied to the estimand.
Not every important task is critical to quality. A formatting error in a nonessential field may require correction but have no effect on safety or interpretation. Treating all errors as equally critical disperses attention and can bury the signals that matter.
Identify factors through cross-functional challenge#
No single function sees the whole trial. Clinicians understand disease course and safety; patients and caregivers identify burdens, meaningful outcomes, and feasibility problems; investigators and coordinators know site workflow; statisticians see threats to estimands, power, and missing-data assumptions; data and technology teams understand source systems and transformations; and laboratories and imaging teams know analytical variation.
A useful design discussion asks:
- What must be true for the result to answer the primary question?
- What failures could harm participants or make the result uninterpretable?
- Where could those failures enter the process?
- Which failures can be prevented by simplifying the design?
- What can be detected early, and what evidence would trigger action?
The discussion should create a reasoned set of factors, not an unfiltered risk register. If everything is labeled critical, the category has lost its function.
Design out avoidable failure#
Prevention is often stronger than detection. An unnecessary visit cannot be missed if it is removed. An ambiguous endpoint cannot be fixed by more source-data review if assessors interpret it differently. A complex eligibility rule invites deviations that a simpler clinically justified rule may avoid.
Examples of preventive design include:
- using a direct, well-defined endpoint instead of a fragile surrogate when feasible;
- reducing nonessential data fields and procedures;
- aligning visits with usual care where this does not compromise the question;
- automating range and temporal checks at data entry for critical variables;
- training with worked endpoint examples and adjudication rules;
- testing randomization, drug supply, and electronic systems before first enrollment;
- piloting participant materials for comprehension and burden;
- specifying how intercurrent events such as treatment discontinuation affect the estimand;
- designing backup procedures for system or laboratory failure.
Complexity should earn its place. Each extra procedure can add burden, deviations, missing data, and site confusion. Simplification is not reduced rigor when it preserves what the study needs to answer.
Assess risk in context#
After identifying a critical factor, the team evaluates what could go wrong and how likely it is. It then evaluates how detectable it would be and the consequence. The goal is not a universal numerical score. It is a transparent basis for choosing controls.
Consider primary-endpoint missingness. Its impact depends on event rate, follow-up duration, and reasons for loss. It also depends on treatment differences and analysis assumptions. A small amount of random missingness may have little effect. Differential missingness caused by adverse effects can seriously bias the treatment comparison.
Controls can operate at several levels:
- protocol design and participant support to prevent loss;
- real-time site metrics to detect emerging imbalance;
- central review of reasons and timing;
- escalation thresholds;
- retrieval of outcomes after treatment discontinuation when consent permits;
- prespecified sensitivity analyses for plausible missing-data mechanisms.
The control plan should match the risk. Monitoring every low-value field at every site can consume resources without improving protection or inference.
Distinguish errors that matter from harmless variation#
Clinical trials contain variation. Visits occur within windows, devices have measurement error, and sites differ in workflow. Quality by design does not demand identical operation in every detail. It asks whether variation threatens a critical factor.
An error's importance depends on direction, concentration, and relation to treatment. A few random late visits may reduce precision. Systematically late endpoint assessment in one randomized group can create bias. One miscoded secondary field differs from incorrect treatment assignment or unblinding.
Aggregate patterns can be more important than individual deviations. A site with modest errors across many participants may threaten results more than one isolated serious correction. Central statistical monitoring can identify unusual distributions, digit patterns, and missingness. It can also identify timing and treatment differences. But signals require clinical and operational interpretation.
Connect controls to evidence and action#
A useful quality plan names the metric, data source, and review frequency. It also names the owner, threshold, and response. “Monitor recruitment quality” is too vague. “Review the proportion of randomized participants who fail central eligibility confirmation by site each month, with investigation above the prespecified threshold” is auditable.
Thresholds should not become automatic declarations that a trial has failed. They are prompts for evaluation. A signal may reflect data lag, a genuine population difference, training failure, or misconduct. The response can include correction, retraining, or process change. It can include targeted review or protocol amendment. When protection or interpretability is at risk, it can include pausing an activity.
Documentation should preserve why a factor was chosen, what evidence was reviewed, and how the plan changed, so that you can reconstruct the reasoning a year later. New safety or feasibility information can alter critical factors during conduct. Quality by design is prospective, but not frozen.
Patient input is a quality control#
Participants can identify burdens that technical teams underestimate. A visit schedule that conflicts with work or caregiving can drive missing data. A lengthy questionnaire may yield rushed or incomplete responses. An endpoint may be measurable yet fail to reflect a benefit patients value.
Early patient and caregiver input can improve endpoint relevance, consent materials, and comparator acceptability. It can improve visit timing, retention, and communication of results. This is not decoration around the protocol. It can directly protect critical factors such as recruitment, adherence, outcome completion, and interpretability.
Input should be representative enough to challenge assumptions. One advocate cannot stand in for every age, culture, disease stage, language, or access constraint. The protocol should record how feedback affected decisions, including the suggestions not adopted and why.
Data and technology remain subordinate to purpose#
Electronic health records, wearables, remote assessments, central laboratories, and decentralized procedures can reduce burden or improve measurement, though they also add dependencies: device adherence, clock synchronization, algorithm versions, data transfer, missing streams, cybersecurity, and vendor change.
For each critical variable, teams should map origin, transformation, and transfer. They should also map review, correction, and analysis. A validated platform does not guarantee that a particular endpoint is valid. The measurement must still represent the clinical concept, operate in the target population, and remain available at the required times.
Source-data verification is only one control. Edit checks, audit trails, and role permissions may be more important for particular risks. So may reconciliation, statistical monitoring, and endpoint adjudication. The mix should follow the critical factor rather than habit.
E8(R1) and E6(R3) now form a current pair#
ICH adopted E8(R1) at Step 4 in October 2021, and the FDA issued final guidance in April 2022. ICH E6(R3) principles and Annex 1 were adopted at Step 4 in January 2025; the FDA issued its final E6(R3) guidance in September 2025.
E8(R1) provides the broad design and development framework. E6(R3) translates related principles into good clinical practice for interventional trials, including participant protection, reliable results, and proportionate approaches. It also includes data governance, roles, and trial conduct. They should be read with topic-specific ICH guidelines rather than as isolated checklists. Guidances also operate through regional adoption, so if your trial is global you have to identify the applicable law, regulation, ethics requirements, and local implementation rather than treating one ICH document as self-executing law everywhere.
A quality-by-design workshop output#
A concise, useful output can give you:
- the primary question, estimand, and decision supported;
- the critical-to-quality factors and rationale;
- principal risks to each factor;
- preventive design features;
- metrics, thresholds, review timing, and owners;
- escalation and remediation paths;
- remaining uncertainty accepted and why;
- links to the protocol, monitoring plan, data plan, safety plan, and analysis plan;
- review points when new evidence could change the assessment.
The artifact should guide action. A large risk table that no one uses is not evidence that quality was designed in.
Sources and further reading
- ICH, E8(R1) General Considerations for Clinical Studies, Step 4 Guideline (2021)
- FDA, E8(R1) General Considerations for Clinical Studies, Final Guidance (2022)
- European Medicines Agency, ICH E8 General Considerations for Clinical Studies
- ICH, E6(R3) Good Clinical Practice, Step 4 Guideline (2025)
- FDA, E6(R3) Good Clinical Practice, Final Guidance (2025)
Questions and answers
Does quality by design mean less monitoring?
Not automatically. It means monitoring is targeted and proportionate to important risks. Some trials may need intensive controls for a critical endpoint or safety issue and lighter controls elsewhere.
Is every protocol deviation a quality failure?
No. Its importance depends on effect on participant protection, critical data, bias, and interpretability. Patterns and causes matter as much as counts.
Can audits create quality after a trial ends?
Audits can detect problems and support correction, but they cannot repair an irrelevant endpoint, an infeasible schedule, or systematic loss of the primary outcome.
Is a critical-to-quality factor the same as a key performance indicator?
No. The factor is the study attribute that must be protected. A performance or risk indicator is a measurement used to watch that factor or its process.
Does E8(R1) prescribe one risk-scoring method?
No. It sets principles. Sponsors and trial teams choose methods suited to the study, document the rationale, and comply with applicable regional requirements.