Evidence explainer

Physician-scientist and medical humanities

Why Promising Discoveries Do Not Reach Patients

A credible laboratory finding is not yet a treatment. Translation is a chain of evidence, engineering, money, and delivery, and any link can give a good reason to stop.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. Discovery and development produce different outputs
  2. Reproducibility is necessary but not sufficient
  3. Candidate engineering changes the problem
  4. The funding chasm
  5. The handoff between research and development
  6. Clinical testing is an attrition process
  7. Regulation is one part of readiness
  8. Bridges that address the middle
  9. Translation is iterative, not a straight pipeline
  10. The later research-to-practice gap
  11. How to read a "breakthrough" claim
  12. References

A laboratory result can be correct and still never affect care. It may identify a disease mechanism but not a safe way to alter it. A candidate may work in a model that poorly predicts human biology. A promising intervention may be too difficult to manufacture, too costly to test, inferior to current care, or impossible to deliver consistently.

The phrase "valley of death" describes parts of this attrition, especially the funding and capability gap between academic discovery and development-ready work. The metaphor is useful when it identifies a missing bridge. It becomes misleading when every stopped project is treated as a tragedy, and some projects should stop because later evidence shows weak benefit, unacceptable harm, poor measurement, or no feasible path to use.

Discovery and development produce different outputs#

Discovery research asks what is true about biology, disease, or behavior. It may reveal a pathway, biomarker, molecular target, imaging pattern, or care-process problem. The output is knowledge supported by experiments and analysis.

Development asks whether that knowledge can be transformed into a reproducible intervention that meets a defined need. Its outputs include a candidate specification, validated assays, and manufacturing methods. They include toxicology, quality controls, and clinical protocols. They include regulatory documentation and evidence comparing benefits with harms.

These activities overlap, but their success criteria differ. A paper can establish an important mechanism without producing a candidate, and a development program may contribute little new basic biology while solving the engineering and evidence problems required for safe use. Treating one as an automatic continuation of the other underestimates the work in between. The FDA's drug-development overview separates discovery and development, preclinical research, clinical research, review, and postmarket safety monitoring, and while other products follow different legal and technical routes, the general lesson travels: a plausible idea must become a defined, testable, consistently produced intervention before a benefit-risk decision is possible.

Reproducibility is necessary but not sufficient#

An early result can fail to replicate because of random error, selective reporting, or small samples. It can fail because of flexible analysis, reagent differences, model instability, or hidden technical conditions. Replication using prespecified methods and relevant models can prevent resources from following a fragile signal.

A replicated finding can still lack translational validity. A model may represent one narrow part of a human disease. The chosen endpoint may not predict clinical benefit. A target can be biologically central but impossible to modify without affecting other essential processes. A biomarker may distinguish groups yet fail to guide a treatment decision.

Translation therefore asks a sequence of harder questions. Is the finding real? Is it causal? Is it modifiable? Can a candidate alter it selectively? Does the model predict human response? Is there a measurable clinical outcome? Can benefit be large enough to justify harm, burden, and cost? Negative answers to any of those are information rather than failure of the research, and a well-designed stop can protect participants and redirect resources sooner.

Candidate engineering changes the problem#

Moving from a target to a candidate involves optimization for potency, selectivity, and stability. It involves optimization for distribution, metabolism, and practical delivery. A diagnostic requires analytic validity, specimen stability, thresholds, and evidence that results change useful decisions. A software or service intervention requires a stable intended use, human-factors work, version control, data governance, and performance evaluation in the intended setting.

These tasks can reveal conflicts. A change that improves one property may worsen another. A compound can reach the target but create unacceptable toxicity. An assay can be accurate in a research laboratory but too variable across routine sites; a model can perform well in its development data yet degrade when workflows, populations, and coding differ.

Manufacturing is part of the scientific claim; if an intervention cannot be produced consistently, stored, tested, and scaled, the result from one carefully made batch does not define the product that patients would receive. Quality systems and process controls are therefore evidence infrastructure, not administrative decoration.

The funding chasm#

Basic-research funding often rewards novel knowledge, training, and publication. Early development needs sustained work on validation, optimization, and manufacturing. It needs work on toxicology, regulatory strategy, and project management. That work can be expensive, specific to one candidate, and less likely to yield a conventional academic publication.

Later investors or development organizations usually seek a clearer candidate, ownership position, and reproducible data package. They seek a regulatory route, a market or public-health need, and milestones that make risk assessable. A discovery can fall between those funding logics: too applied for another discovery award, but too incomplete for development capital.

This is the core of the valley-of-death metaphor described in translational policy literature. It is not simply a shortage of total research spending. Timing, allowable costs, expertise, incentive structures, and who can own the next decision all matter. The gap can run deeper for rare diseases, noncommercial public-health interventions, and needs concentrated in populations with limited purchasing power, because expected financial return may not match social value. Public, philanthropic, nonprofit, and shared-development mechanisms can play different roles there, but each still needs clear evidence standards and governance.

The handoff between research and development#

A publication is rarely a complete handoff package. If you are the receiving team, you may need protocols, raw and processed data, code, materials, assay performance, failed experiments, batch records, model limitations, intellectual-property status, and permission to use essential resources. If key knowledge exists only in one laboratory member's memory, transfer is fragile.

The scientific question also changes. A research team may have optimized an experiment to detect a mechanism. A development team needs a method robust to different operators, sites, batches, and time. Repeating the same experiment is not enough; the method must survive variation expected in its future use.

Roles can be ambiguous. Who owns replication? Who pays for process development? Who chooses a clinical indication? Who can stop the program? Who communicates with regulators and patient communities? Delayed answers create a period in which everyone supports the idea but no one owns the next deliverable. A strong handoff answers them on paper: the intended product or practice, the unmet need, the target population, the evidence gaps, the decision milestones, the responsibilities, the data rights, and the criteria for stopping. It also preserves the negative data, because a hidden failure is one the next team repeats.

Clinical testing is an attrition process#

Human studies ask questions that preclinical work cannot settle. Early studies characterize safety, behavior in the body, and feasible administration. Later studies examine efficacy and safety in defined populations and comparisons. Designs vary by product and condition, and development phases are not guarantees of progress.

Analyses of clinical-development programs show substantial attrition and differences across therapeutic areas and phases. Exact success estimates depend on the database, cohort dates, definitions, and whether programs with unknown outcomes are included; a single headline percentage is not a universal constant you can carry between fields.

Clinical failure can reflect inadequate benefit, harm, or a wrong dose strategy. It can reflect weak target biology, poor endpoint selection, operational problems, or an effect limited to a subgroup not identified in advance. A subgroup story found after many analyses needs confirmation; it cannot automatically rescue a failed primary question. Ethical development builds stop rules and monitoring into the program. Continuing because prior spending was large is a sunk-cost error, not commitment to patients.

Regulation is one part of readiness#

Regulators evaluate evidence under the legal framework for the product and intended use. For new drugs, the FDA reviews clinical, nonclinical, and manufacturing information. It also reviews labeling and other submitted information, then decides whether benefits outweigh risks for the proposed use. Devices, biologics, diagnostics, and other categories have distinct pathways.

Early regulatory planning can tell you whether the evidence you plan to collect will answer the necessary questions, and it can align assay validation, manufacturing controls, trial endpoints, population definitions, and safety monitoring before an expensive study begins.

Regulatory clearance or approval does not mean risk is zero, benefit is universal, or comparative value is established against every alternative. It means the applicable standard was met for a defined product and use based on the submitted evidence, and postmarket monitoring continues because larger and more varied use can reveal new information.

Framing regulation as the sole barrier misses earlier failures in biology, reproducibility, engineering, and funding. Framing it as a final proof misses the work of implementation and continuing safety assessment.

Bridges that address the middle#

Translational cores can supply shared capabilities that a single laboratory cannot maintain, such as medicinal chemistry, toxicology planning, and biostatistics. Others are assay validation, manufacturing consultation, informatics, and regulatory support. Shared infrastructure reduces the need to rebuild expertise for every project.

Milestone-driven funding connects resources to defined, decision-relevant deliverables. The NCATS Translational Research in Neglected Diseases program describes a team-based model with a gap analysis, project plan, timelines, deliverables, and go or no-go points. The value of this structure is not that every project advances. It is that continuation and stopping are linked to evidence agreed in advance.

Team science brings discovery researchers, developers, and clinicians into the same problem definition. It brings statisticians, engineers, and regulatory specialists. It brings manufacturing experts, implementation researchers, and patient or community partners. NCATS identifies partnerships and cross-disciplinary work as central to translational science.

Partnership agreements can clarify data, materials, and publication. They can clarify intellectual property, decision rights, and resource commitments. Technology-transfer offices can support licensing and contracts, but transfer succeeds only when the technical package is sound.

Patient and community input can refine which outcomes matter, whether procedures are acceptable, and what barriers would prevent use. That input is not a substitute for controlled evidence. It helps ensure that development answers a relevant question and that later delivery is plausible.

Translation is iterative, not a straight pipeline#

The familiar arrow from discovery to preclinical work to trials to practice is a map, not a law. Clinical observations can generate laboratory questions. Manufacturing can force candidate redesign. Early human data can change a target hypothesis. Implementation can teach you that an outcome or workflow assumption was wrong all along.

Iteration is productive when changes are documented and tested, and it becomes risky when teams rewrite the goal after seeing unfavorable results, hide failed attempts, or carry a candidate forward without a decision standard. A learning translational system preserves traceability instead: which hypothesis was tested, what changed, why it changed, and what evidence supports the next version, which is the record that makes both the science and the handoff possible.

The later research-to-practice gap#

Even an approved and available intervention may not improve population health. Clinicians need to know which patients are eligible, how the option compares with current care, how to monitor it, and whether the health system can deliver it. Patients face cost, distance, and language. They face trust, time, disability access, and competing responsibilities.

Implementation science studies methods for integrating evidence-based practices into routine settings. It examines reach, adoption, and fidelity. It examines adaptation, sustainability, and context. A trial result obtained under intensive support may not reproduce in your service when staffing, reminders, or follow-up differ.

This later gap needs its own evidence and budget. Training, workflow redesign, data systems, quality monitoring, and equitable access are not automatic consequences of approval. Nor should adoption be pursued when comparative benefit, feasibility, or affordability is weak.

The endpoint of translation is not a product on a shelf or a guideline sentence. It is reliable benefit in the people and settings the intervention is intended to serve, with harms and unequal outcomes monitored.

How to read a "breakthrough" claim#

Ask where the work actually sits. Is it a mechanism, model, or candidate? Is it a preclinical study, early human study, or comparative trial? Is it a regulatory decision or an implementation result? Each supports different conclusions.

Check what has been replicated, whether outcomes are clinical or surrogate, and whether the intervention is defined well enough to reproduce. Look for the next unresolved dependency: manufacturing, toxicity, or trial recruitment. It could equally be comparator, regulatory evidence, cost, or delivery.

Then ask who owns the next milestone and what result would stop the program. A claim that names no remaining uncertainty is usually describing aspiration to you, not readiness.

Many discoveries deserve further work. Some deserve to stop. A strong translational system makes both decisions earlier, more transparently, and with the eventual patient need in view.

References#

  1. Mapping the translational science policy valley of death
  2. Clinical development success rates by phase and indication
  3. FDA drug development process
  4. NCATS overview and translational science mission
  5. NCATS Translational Research in Neglected Diseases operational model
  6. NCATS Clinical and Translational Science Awards Program
  7. Research-to-practice gap and implementation science

Questions and answers

What is the translational valley of death?

It is the gap in funding, capabilities, and ownership between a research discovery and development-ready work. The term is most useful when it identifies the specific bridge that is missing.

Does a discovery that never reaches patients mean the science was wrong?

Not always. The finding may be valid but not safely modifiable, manufacturable, clinically useful, financeable, or deliverable. Later evidence can also reveal a sound reason to stop.

What makes a research-to-development handoff work?

A reconstructable package includes methods, data, materials, negative findings, validation, ownership, intended use, evidence gaps, milestones, roles, and stop criteria. A publication alone rarely supplies all of it.

How can milestone funding help?

It links continued resources to prespecified technical and evidence goals. This supports focused work and timely stop decisions rather than assuming every funded project must advance.

Is approval the end of translation?

No. Adoption, access, workflow, training, comparative use, monitoring, and sustainability determine whether an available intervention produces benefit in routine settings.