Evidence explainer

Health policy, systems, and equity

Why Pilots Do Not Scale

Pilots run on selected sites, committed teams, extra money, and somebody fixing things fast. Scale changes all four. What a pilot really produces is a tested set of assumptions.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. A pilot has a narrower job
  2. A favorable p-value does not validate a pilot
  3. Progression criteria make decisions auditable
  4. Pilots often select unusually favorable settings
  5. Champions can conceal the labor model
  6. Temporary funding distorts the cost picture
  7. The denominator changes at scale
  8. Fidelity and adaptation are not opposites
  9. Training decay is predictable
  10. Supply chains become part of effectiveness
  11. Governance determines whether learning continues
  12. Equity can worsen during expansion
  13. Scale can change the effect itself
  14. Implementation outcomes explain the result
  15. Replication should seek contrast, not comfort
  16. What a credible scale decision includes

A pilot can demonstrate that a team recruited 80 people, delivered a new service, collected the planned data, and solved problems for six months, but it cannot, by itself, show that the same service will work across 500 sites, fit ordinary budgets, reach underserved groups, or survive after the founding team leaves.

The distinction is easy to lose, because "pilot successful" lands on you as "intervention proven." In research, a pilot or feasibility study usually asks whether a larger evaluation can be done and how it should be designed. In health systems, a demonstration may also test delivery. Neither purpose turns a small favorable result into evidence for population-wide impact.

WHO's 2026 scaling guidance treats scale as a structured process of exploring, adapting, learning, and embedding an innovation in public systems. It is not photocopying a pilot. The intervention and the system change each other. So scale requires evidence about context, ownership, and capacity. It requires evidence about finance, equity, and sustained performance.

A pilot has a narrower job#

The CONSORT extension for pilot and feasibility trials emphasizes that their objectives concern feasibility. Can eligible participants be found? Will they consent? Can the intervention be delivered as designed? Are outcome measures completed? Can randomization and follow-up operate? What variance or event-rate assumptions should inform the larger trial?

Those are legitimate and consequential questions. A full trial can waste money and place participants at risk if recruitment is unrealistic or the outcome cannot be measured; a pilot can reveal that the intervention needs redesign before effectiveness is tested. The target is not a small version of every later conclusion, and a feasibility study needs methods and analyses matched to feasibility outcomes, not a conventional p-value attached to an underpowered clinical endpoint.

A favorable p-value does not validate a pilot#

Small studies produce imprecise estimates. A statistically nonsignificant clinical result may reflect inadequate information, not lack of benefit. A statistically significant result may be an unstable overestimate selected from many outcomes or subgroups.

Effect sizes from small pilots can fluctuate sharply because a few participants have large influence, so using the observed pilot effect as the exact assumption for a definitive trial can lead to an undersized study. Planning should consider clinically important differences, uncertainty, prior evidence, and conservative nuisance-parameter estimates. Clinical outcomes can still be collected to test measurement and look for safety concerns, but they should be reported with confidence intervals and an explicit statement that the study was not designed to establish effectiveness.

Progression criteria make decisions auditable#

Before results are known, a pilot should define thresholds that guide proceed, amend, or stop decisions. Criteria may cover recruitment rate, retention, and missing data. They may cover delivery fidelity, acceptability, and serious safety issues. They may cover time and resource use.

A traffic-light approach can be useful: green supports moving forward, amber requires modification, and red indicates a major barrier. Thresholds need interpretation rather than mechanical obedience. Missing a target by a small amount may be repairable, while meeting recruitment through an unsustainable advertising effort may reveal a deeper problem. Publishing the criteria and the reasoning behind deviations reduces hindsight bias. It also gives funders, communities, and later implementers a clear record of what was learned.

Pilots often select unusually favorable settings#

Early sites may be chosen because they have strong leadership, compatible infrastructure, and staff who asked to participate. Those features help discover whether the idea can work. They also limit transfer to sites with vacancies, competing priorities, poor connectivity, or different patient needs.

Participants can be similarly selected. Strict eligibility, extensive consent conversations, transport support, or digital access requirements may create a group unlike the population a national service is expected to reach. Scale is where all that variety arrives at once. That is why performance has to be studied in settings that differ in geography, workload, and language. Those settings also differ in resources and baseline outcomes.

Champions can conceal the labor model#

A motivated founder may answer messages after hours, retrain staff personally, negotiate missing supplies, and repair data errors. Researchers may remind participants, coordinate referrals, and prepare reports. The program appears efficient because much of its labor sits outside the budget.

Routine scale requires named roles, realistic caseloads, supervision, backup, and compensation. Tasks that depend on exceptional effort need redesign or explicit funding. Burnout is not a scale strategy.

Time-and-motion data, workflow mapping, and process logs can reveal hidden work, because the question you want answered is not only whether delivery occurred, but who made it occur, and what else that person could not do while they were doing it.

Temporary funding distorts the cost picture#

Pilot budgets may cover devices, training, travel, data systems, and project management through a grant. A health service later has to place those costs inside recurring budgets and procurement rules.

Economies of scale can lower the unit cost of software, training materials, or centralized purchasing. Diseconomies can raise it through regional coordination, supervision, and help desks. They can raise it through quality assurance, replacement stock, translation, and service for remote areas.

Cost per enrolled participant is also sensitive to reach. A pilot that serves people who are easiest to contact may look inexpensive. Reaching the final groups necessary for equitable coverage may require more resources, and that added cost may be justified.

The denominator changes at scale#

A pilot report may describe outcomes among people who enrolled and completed the program. A system decision concerns everyone eligible. The gap includes people never invited, unable to access the service, or excluded by technology. It includes people lost during referral or unwilling to continue.

So measure reach against a defined source population. Ask for adoption among sites and staff, not just patients. Then ask whether use and benefit persisted once the launch support was withdrawn. An intervention can retain a strong effect among participants while having little population impact because only a small or advantaged fraction receives it.

Fidelity and adaptation are not opposites#

Complex interventions usually contain core functions and adaptable forms. A reminder system's function may be timely prompting; its form could be text, call, paper, or an in-person cue. Insisting on one form can make implementation brittle. Changing the core function can remove the mechanism.

The Medical Research Council framework asks researchers to consider context, program theory, and stakeholders. It asks them to consider uncertainty, refinement, and economic factors across development and evaluation. Adaptation should be documented with its rationale and effect, not treated as protocol failure by default. Scale teams need a clear theory of how activities produce outcomes. That theory helps decide what must remain stable and what should change for local fit.

Training decay is predictable#

Pilot staff often receive intensive initial training from the designers. At scale, staff turnover, schedule conflicts, and varying baseline skills dilute that model. A one-time workshop cannot carry a multiyear service.

A training system that survives needs competency checks, refresher support, and supervision. It needs accessible materials and a plan for the people hired next year. And what you measure has to show whether the intended practices are still being used, not how many people signed the attendance sheet. If specialized skill is scarce, the program may need task redesign, decision support, regional consultation, or a slower rollout. The workforce plan is part of the intervention.

Supply chains become part of effectiveness#

A clinical or public-health service cannot work when diagnostic supplies, medicines, devices, forms, or connectivity fail. Small pilots can keep reserve stock and use personal contacts to avoid delays. National procurement moves through forecasting, contracts, and warehousing. It moves through distribution, maintenance, and quality control.

Stockouts and device downtime should be treated as implementation outcomes because they alter the intervention received. Redundancy, repair, inventory data, and vendor accountability belong in the scale design. The same applies to referral capacity. Screening more people without confirming diagnosis or providing treatment can create queues rather than health benefit.

Governance determines whether learning continues#

At pilot scale, decisions may sit with one principal investigator or program director. At larger scale, ministries, payers, and professional bodies have different authority and incentives. So do local leaders, vendors, and communities.

Governance needs a responsible owner, decision rights, and data stewardship. It needs safety escalation, procurement accountability, and a method for changing the program. Without those structures, local teams improvise inconsistently or wait for a central group that cannot respond quickly. WHO's scaling framework places government and system stewardship at the center. Long-term adoption is more plausible when the receiving system helps shape priorities, design, evidence needs, and financing from the beginning.

Equity can worsen during expansion#

Average uptake can rise while gaps widen. Digital programs may favor people with reliable devices, literacy, privacy, and connectivity. Clinic programs may favor people who can travel during working hours. Language, disability, legal status, stigma, and prior mistreatment can affect access and trust.

Pilots should disaggregate reach, retention, outcome, and burden by relevant groups while protecting privacy. Qualitative research can identify barriers that aggregate metrics hide. Co-design and sustained public participation can reveal whether the proposed service solves the problem people actually face. Equity is not an optional analysis after expansion. It shapes eligibility, channels, staffing, safeguards, and resources.

Scale can change the effect itself#

Treatment effects are not portable constants. A program delivered by specialists may have a different effect when delivered by general staff. Usual care differs across settings, so the added value of an intervention changes. Broader eligibility can alter baseline risk and room for improvement.

Contamination may rise when trained staff serve both intervention and comparator groups. Network effects can strengthen a public-health program, while congestion can weaken it. Policy changes during rollout can also modify outcomes. Pragmatic trials, staged rollouts, and cluster designs can study performance under routine conditions. So can interrupted time series and other approaches. Design choice should follow the causal question and operational constraints.

Implementation outcomes explain the result#

Effectiveness tells whether outcomes changed. Implementation outcomes help explain why. Common constructs include acceptability, adoption, and appropriateness. They include feasibility, fidelity, and cost. They include reach and sustainability.

These measures need precise definitions. "High adoption" is meaningless without a denominator and timeframe. "Good fidelity" requires observable components and a scoring rule. Mixed methods can connect numeric trends to the decisions, constraints, and experiences behind them. Process evaluation is especially important when the main outcome is disappointing. It distinguishes a weak theory from failed delivery, although post hoc explanations still require caution.

Replication should seek contrast, not comfort#

Repeating a pilot in a nearly identical favorable site tells you less than running it somewhere with different constraints. Deliberate contrast is what reveals which contextual factors are essential and which can be adapted.

Replication does not mean freezing every detail. It means making the intervention, context, changes, and outcomes sufficiently transparent to compare. The TIDieR approach and related reporting guidance help specify materials, procedures, and providers. They help specify dose, setting, and modifications. A sequence of replications can reduce uncertainty before broad expansion. It may also identify that the innovation works only for a narrower population or setting, which is a useful result.

What a credible scale decision includes#

If you are the one writing the decision memo, state the problem, the target population, the intervention theory, the strength of the evidence, the feasibility findings, the expected benefit, the harms, the total cost, the equity effects, the workforce and supply requirements, the governance, and the alternatives. Then mark which of those are facts and which are assumptions.

The rollout can include explicit learning phases, stop rules, safety monitoring, and an evaluation design. Expansion speed should match the reversibility of harm and the ability to detect poor performance. A service can be phased without pretending that every phase is still only a pilot.

Success at scale means more than activity. It means sustained, equitable improvement relative to a credible alternative at an acceptable opportunity cost.

Sources and further reading

  1. WHO guidance and toolkit for scaling innovations in public health systems, 2026
  2. WHO practical guidance for scaling health-service innovations, 2009
  3. CONSORT extension for randomized pilot and feasibility trials
  4. Medical Research Council framework for developing and evaluating complex interventions, 2021
  5. Implementation outcomes framework

Questions and answers

Is a pilot just a small randomized trial?

Not necessarily. A randomized pilot tests whether trial procedures are feasible. Other pilots may test service delivery without randomization. In both cases, objectives should be explicit.

Can pilot clinical outcomes be reported?

Yes, but they should be presented as imprecise and secondary to feasibility unless the study was actually designed and powered for effectiveness.

What should happen after a successful pilot?

The team should review progression criteria, repair identified problems, test uncertain assumptions, and choose a justified next evaluation or rollout phase.

Does adapting a program destroy fidelity?

Not if core functions are preserved and changes are documented and evaluated. Adaptation can improve fit; untracked changes can make the result uninterpretable.

Why can cost per person rise at scale?

Expansion adds coordination, supervision, remote delivery, quality assurance, replacement, and equitable-reach costs that a selected pilot may not include.