Evidence explainer

Physician-scientist and medical humanities

The Translation Gap Is Not One Gap

A discovery can fail to improve health at many handoffs: model to human, trial to decision, recommendation to workflow, and service to equitable reach. Each handoff needs a different method.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Discovery is not yet a candidate intervention
  2. Preclinical evidence can be technically correct and clinically unhelpful
  3. Human testing changes the question again
  4. Approval is not identical to implementation
  5. Implementation science studies the delivery problem
  6. Reach changes the population effect
  7. Adaptation and fidelity are not opposites
  8. Deimplementation is part of translation
  9. Patient and community knowledge changes upstream design
  10. Speed comes from reducing rework
  11. Build an evidence chain for the intended benefit
  12. References

“Bench to bedside” suggests a straight trip from a laboratory discovery to a patient benefit. The path is neither straight nor one-directional. A molecular finding may not reproduce. An animal model may not predict human biology. A successful trial may compare the wrong alternatives, use an outcome that matters little in practice, or enroll a narrow population. A useful intervention may then fail because it is unaffordable, difficult to deliver, or incompatible with workflow.

The translation gap is therefore a set of linked gaps. Treating it as one delay encourages one cure, usually more speed or more funding. Each handoff instead requires its own question, design, expertise, and evidence of success.

Discovery is not yet a candidate intervention#

A laboratory association may identify a pathway, biomarker, or target. Translation first asks whether the finding is reproducible, causal, and relevant to human disease. Repeating the experiment in the same model is useful but limited. Orthogonal methods, different laboratories, human tissue, genetic evidence, and perturbation can test whether the mechanism survives changes in technique.

Model choice matters. Cell lines can accumulate changes and lack organ context. Organoids recreate selected architecture but not full circulation, immunity, or behavior. Animal models can isolate mechanisms yet differ in metabolism, lifespan, anatomy, and disease induction; a model should be judged by the feature it represents, not by whether it resembles the entire human condition.

A target also needs tractability. Can it be changed safely? Is the relevant tissue reachable? Does intervention occur before irreversible disease? Is there a biomarker showing that the proposed mechanism was altered? These questions connect basic and preclinical work.

Preclinical evidence can be technically correct and clinically unhelpful#

An intervention may produce a large effect under tightly controlled conditions using young, genetically similar animals, treatment before disease, and an outcome measured shortly afterward. Human use may involve older people, comorbidity, established disease, different dosing, and concurrent medicines. None of that was in the experiment.

Randomization, blinding, sample-size planning, prespecified outcomes, and complete reporting strengthen preclinical studies. Replication across sex, age, relevant disease stage, and model can reveal boundary conditions. Negative findings should remain visible so later teams do not repeat a dead end.

More experiments are not automatically better. A weak target can accumulate publications because each model confirms a nearby question. Translation benefits from explicit milestones and stopping rules: what result would justify advancing, redesigning, or ending the program?

Human testing changes the question again#

Early clinical studies examine safety, dose, pharmacokinetics, pharmacodynamics, and proof of mechanism. A change in a biomarker can show that the intervention reached its target. It does not necessarily show that people feel better, function better, or live longer.

Later trials estimate benefit and harm under a defined protocol. Their relevance depends on comparator, eligibility, and outcome. It depends on follow-up, adherence, and setting. A placebo comparison may establish specific efficacy while leaving uncertainty about the best existing alternative, and a composite endpoint may increase event counts while being driven by a less important component.

External validity is designed, not added later. Broad eligibility can improve applicability, but heterogeneity may require larger samples and planned subgroup questions. Pragmatic elements can embed recruitment, treatment, and outcome collection in routine care.

Approval is not identical to implementation#

Regulatory approval answers whether a product meets the applicable standard for a defined use. A health system still must decide who should receive it, how it fits existing care, what infrastructure is needed, how benefit and harm will be monitored, and whether the opportunity cost is acceptable.

Guidelines synthesize evidence and make recommendations. A recommendation can remain unused because clinicians do not know it, disagree with it, lack time, cannot order the intervention, or work in a payment system that rewards another action. Patients may face travel, language, cost, disability, or trust barriers. The intervention delivered in practice is often a package: drug or procedure, diagnostic pathway, and staff training. The package also holds scheduling, documentation, supply chain, payment, and follow-up. Failure of any component can erase the trial's benefit.

Implementation science studies the delivery problem#

Implementation science asks which strategies help evidence-based interventions become routine, equitable, and sustainable. It distinguishes the clinical intervention from the implementation strategy. A clinical intervention might be a vaccination program. Strategies might include standing orders, reminders, outreach, staff training, or mobile delivery.

The updated Consolidated Framework for Implementation Research, or CFIR, organizes determinants across five domains. They are innovation, outer setting, inner setting, individuals, and implementation process. It is a diagnostic framework, not a guaranteed recipe. You select the constructs that are relevant to your setting and test how they affect implementation.

Proctor and colleagues distinguish implementation outcomes from service and clinical outcomes. The implementation outcomes include acceptability, adoption, appropriateness, and feasibility. They include fidelity, cost, penetration, and sustainability. A program can improve adoption without improving health if delivery quality is poor. It can improve health in early adopters yet fail to reach most eligible people.

Reach changes the population effect#

RE-AIM examines reach, effectiveness, adoption, implementation, and maintenance. A highly efficacious intervention with low reach may have less population value than a modest intervention delivered broadly.

Reach is not only a percentage. Compare participants with the intended population by age, health, and language. Compare by geography, income, disability, and other relevant characteristics. Adoption also has levels: which clinics, clinicians, and organizations participate?

Maintenance asks whether benefit and delivery persist after initial funding or champions leave. A pilot can succeed because a research team performs work that ordinary staffing cannot sustain. Cost, turnover, competing priorities, and data burden become part of effectiveness at scale.

Adaptation and fidelity are not opposites#

Fidelity means preserving the functions responsible for benefit. Adaptation changes form or delivery to fit context. Translating materials, changing visit length, or using a different staff role may improve reach without weakening the core mechanism.

The challenge is to specify core functions and adaptable forms. If every component is declared essential, the program may be impossible to implement. If unrestricted changes are allowed, the tested intervention disappears. Document what changed, why, who decided, and what happened to outcomes. Iterative learning can improve fit while preserving an auditable evidence chain.

Deimplementation is part of translation#

Health systems have limited time, staff, and money. Every new practice you add without removing low-value work increases the burden. Deimplementation studies how to reduce care that is ineffective, harmful, or no longer supported.

Stopping can be harder than starting. Existing practices become embedded in order sets, training, quality measures, reimbursement, and professional identity. Evidence of no average benefit may be resisted if clinicians recall vivid successes. Patients may interpret removal as rationing. Replacement strategies can help: pair the recommendation to stop with a safer alternative, revise defaults, monitor unintended consequences, and explain the evidence. Translation should measure what the new practice displaces.

Patient and community knowledge changes upstream design#

NCATS describes patient involvement as important across the translational spectrum; people living with a condition can identify outcomes that matter, burdens a protocol overlooks, and definitions that do not match lived experience. Community partners can identify access routes, historical reasons for mistrust, and communication needs.

Participation should have defined influence. Asking for feedback after the primary endpoint and visit schedule are fixed is consultation, not co-design. Compensation, accessibility, language, and feedback about how input changed the study are part of credible involvement.

Patient priorities do not replace scientific validity. They keep a valid study pointed at a useful question.

Speed comes from reducing rework#

Faster translation does not require skipping safety or accepting weak inference. It can come from reusable platforms, common data standards, and adaptive but prespecified designs. It can come from parallel manufacturing planning, shared controls, rapid contracting, and early regulatory discussion.

It also comes from ending weak projects. Publication incentives favor positive novelty, while translational progress needs reliable negative results and explicit reasons for stopping. Versioned protocols and data standards make findings easier to combine.

Ioannidis argues that useful clinical research should address an important problem, be placed in existing evidence, use patient-centered outcomes, provide value for money, and be feasible in practice. Those criteria turn “translation” from a promise into a design property.

Build an evidence chain for the intended benefit#

For your own innovation, write the chain out: target mechanism, model validity, and human proof of mechanism. Continue with patient-important effect, comparative benefit, and delivery requirements. Finish with implementation outcomes, equity, and population impact. Then mark which links you have measured and which you are assuming.

Assign a method to each uncertainty you marked. Mechanistic experiments cannot answer adoption. A randomized trial cannot alone predict sustainability. Interviews cannot estimate treatment effect. A learning health system can combine methods while keeping their questions distinct.

The related article on population research and clinical practice describes one route from cohorts to decisions, and real-world evidence after approval covers the continuing evidence cycle.

The translation gap closes through aligned handoffs, not a heroic final leap. Better science makes each transition testable and feeds failures back to the stage that can correct them.

Metrics should follow the handoff. Count replication and target validation before claiming a candidate. Measure recruitment, retention, and patient-important outcomes during trials. Measure reach, fidelity, and burden during implementation. Measure cost, equity, and maintenance. Track health outcomes after scale. A dashboard that reports publications and patents alone rewards activity near the start while leaving public benefit assumed, and translation becomes accountable when each promised transition has an observable criterion and a named party responsible for learning from failure.

References#

  1. NCATS translational science spectrum
  2. NCATS translational science principles
  3. RE-AIM planning and evaluation framework
  4. Updated CFIR
  5. Outcomes for implementation research
  6. Why most clinical research is not useful

Questions and answers

What does translational research mean?

It refers to linked research that moves among biological discovery, preclinical work, human studies, clinical implementation, and population health, with feedback in both directions.

Is the translation gap mainly the time needed for drug approval?

No. Regulatory review is one handoff. Evidence can also fail at replication, trial relevance, guideline uptake, workflow integration, affordability, or equitable reach.

How is implementation science different from an efficacy trial?

An efficacy trial asks whether an intervention can work under defined conditions. Implementation science asks how it can be adopted, delivered, sustained, and adapted in real settings.

Does faster translation require lower evidence standards?

No. It requires earlier alignment of questions, reusable methods, parallel planning, transparent uncertainty, and stopping weak projects sooner.

Why involve patients and communities before the final trial?

Early involvement can improve question relevance, outcomes, burden, recruitment, consent, delivery, and interpretation before expensive design choices become fixed.