Evidence explainer

Chronic disease in primary care

Pay for Performance in Primary Care: What the Quality Evidence Actually Shows

Financial incentives can improve some recorded care processes, especially when targets are clear and easy to measure. Evidence for durable health gains, value, and equity is much less certain.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. What pay for performance is trying to fix
  2. Why process measures move first
  3. What systematic reviews have found
  4. The measurement problem inside the apparent result
  5. Thresholds can produce distorted effort
  6. Equity can improve or worsen
  7. The opportunity cost is clinical, not merely administrative
  8. How to appraise a program
  9. Better design starts with restraint

Pay for performance links part of a clinician's or organization's payment to specified measures. In primary care, those measures might count blood-pressure checks, laboratory monitoring, vaccinations, medication reviews, or achievement of a clinical threshold. The appealing theory is straightforward: reward measurable quality and quality should rise. The evidence is less tidy.

Financial incentives often improve the activities that are directly counted, particularly documentation and relatively simple care processes. Effects on outcomes that patients can feel, such as fewer cardiovascular events, better function, or longer life, are smaller, inconsistent, or not well measured. Programs can also redirect attention, widen or narrow inequities, and create large administrative costs. A fair judgment therefore asks more than whether a performance score increased.

What pay for performance is trying to fix#

Primary care contains many evidence-based actions that are easy to postpone during a crowded visit, and a registry reminder plus a financial incentive may make a clinic more systematic about finding overdue monitoring or recalling eligible patients. At an organizational level, payment can fund data staff, team-based workflows, and population management that ordinary fee structures do not support.

This logic is strongest when the desired action is well supported, within the practice's control, reliably measurable, and important enough to justify the effort. It becomes weaker when a measure is a crude proxy for quality, depends heavily on social circumstances, or encourages a single disease target at the expense of a patient's larger situation.

The incentive is also only one part of the intervention. Public reporting, electronic prompts, clinical education, staffing, audit feedback, and national quality campaigns often arrive at the same time. You cannot automatically assign an observed change after launch to the payment component.

Why process measures move first#

A process measure asks whether an action occurred. Was kidney function checked? Was smoking status recorded? Was an indicated review completed? These measures offer a clear task, a defined denominator, and a short path between effort and score. They are therefore responsive to reminders and payment.

An outcome measure sits farther downstream. Blood pressure, glycated hemoglobin, hospital admission, or survival reflects treatment, biology, adherence, continuity, deprivation, competing illness, and chance. Some outcomes take years to change. Others can improve on paper if the practice attracts a healthier case mix or excludes people through permitted exception rules.

That difference does not make process measures trivial. A well-chosen process can represent a necessary link in effective care. It does mean that a six-point rise in a process score should be described as a six-point rise in that process, not as proof of an equivalent health gain.

What systematic reviews have found#

The 2022 Cochrane review found that the evidence base was too limited and uncertain to support a broad conclusion for or against financial incentives. Several included studies reported modest improvement in selected professional behaviors, but study designs, programs, measures, and risk of bias varied. Patient outcomes were much less certain.

A 2025 BMJ systematic review of primary-care pay-for-performance programs similarly found improvement in recorded quality after introduction, especially early and for measured processes, and the pattern was not uniform across measures or over longer follow-up. Importantly, its withdrawal analyses did not show that gains simply persisted untouched: removal of incentives was often followed by deterioration in recorded quality, with larger declines for some more complex processes.

The lesson is not that incentives always work or that withdrawing them always reverses care. It is that behavior can remain linked to the surrounding measurement and payment system. A process that has become genuinely embedded may persist; one sustained by reminders, staff time, and financial priority may fade when those supports disappear. Recent review of locally designed programs in UK primary care points the same way, toward improvements in selected processes alongside inconsistent evidence for clinical outcomes and major gaps in patient-reported outcomes and cost effectiveness. Different programs should not be treated as a single drug with one stable effect.

The measurement problem inside the apparent result#

Performance data can improve through several pathways:

  1. More eligible patients receive better care.
  2. Care that was already occurring is documented more completely.
  3. Coding becomes more aligned with the measure specification.
  4. Follow-up becomes concentrated near the scoring threshold.
  5. Denominator or exception rules alter who is counted.

The first is the intended pathway. The second can still be useful because reliable records support continuity and safety. The remaining pathways range from defensible data correction to strategic behavior that weakens the measure's meaning.

Evaluators should compare clinical records, claims, patient outcomes, and audit samples where possible. They should report changes in denominator composition, missingness, exclusions, and coding intensity. A score handed to you without its data-generating process is difficult to interpret.

Thresholds can produce distorted effort#

Suppose payment begins when a practice reaches a target proportion or when an individual result crosses a cutoff. Effort may cluster around people who are closest to the threshold, because they offer the largest expected score change. Patients with severe illness, unstable housing, language barriers, or multiple competing priorities may require more time while contributing no additional points.

This is sometimes called effort diversion or teaching to the test. It does not require fraud or bad motives. It can emerge rationally when a busy team responds to the rules it has been given. Continuous measures, improvement rewards, safeguards for complex patients, and monitoring of untreated need can reduce the pressure, but every design has tradeoffs.

Equity can improve or worsen#

A universal incentive might narrow gaps if it helps practices build reliable recall systems for people previously missed; it might widen gaps if well-resourced clinics can respond faster, if measures do not adjust fairly for need, or if hard-to-reach patients are viewed as threats to performance.

Simple risk adjustment is not a complete answer. Adjusting away poor outcomes associated with deprivation can normalize inequity, while failing to adjust can penalize practices caring for populations with fewer resources. Evaluations should present results by deprivation, race and ethnicity where data quality permits, disability, language, rurality, and baseline risk. They should also examine whether money flows toward or away from communities with greater need. Patient experience belongs in that list too, because a numerically successful program can still worsen access or trust if your visit feels dominated by checklists that do not match what you came in for.

The opportunity cost is clinical, not merely administrative#

Every measure consumes attention. Staff must identify eligible patients, reconcile data, contact people, document exceptions, and answer audits. That work may be worthwhile, but it displaces something else.

Unmeasured areas can include diagnostic reasoning, continuity, coordination, symptom relief, caregiver support, and care for conditions without a convenient metric. Multimorbidity creates the sharpest conflict. Completing every single-disease checklist may produce treatment burden or recommendations that do not fit a person's goals. So a strong program limits the measure set, rotates or retires low-value items, protects time for person-centered care, and evaluates administrative workload, and it does not mistake gross incentive payments for net benefit before implementation and reporting costs have been subtracted.

How to appraise a program#

Start with the causal chain. What exact behavior is being paid for, why should it improve health, and how long should that take? Then ask yourself:

Randomized evaluations are possible for new incentives, especially with cluster allocation or phased rollout. When randomization is unavailable, interrupted time-series and controlled difference-in-differences designs are stronger than a single before-and-after comparison. All of them need you to check for preexisting trends, concurrent reforms, and changes in data definitions.

Better design starts with restraint#

No payment system can measure the whole of primary care. A more credible design treats metrics as a limited set of signals rather than a complete definition of quality; it combines evidence-based processes with outcomes, safety measures, patient-reported experience, and equity checks. It rewards meaningful improvement without making baseline excellence a disadvantage.

Programs also need an explicit revision process. Measures can become outdated after guidelines change, reach a ceiling, or create behavior that was not anticipated. Retirement is a sign of governance, not failure. Changes should be announced, versioned, and evaluated so that apparent trends are not artifacts of a rewritten specification.

Sources and further reading

  1. Cochrane Review, Financial Incentives for Healthcare Professional Behaviour and Patient Outcomes (2022)
  2. Minchin and colleagues, Effects of Pay for Performance on the Quality of Primary Care, BMJ (2025)
  3. Mandavia and colleagues, Effects of Pay for Performance on the Quality of Primary Care, BMJ Open Quality (2021)
  4. Systematic Review of Local Pay-for-Performance Programs in UK Primary Care, PubMed (2026)

Questions and answers

Does pay for performance improve primary care?

It can improve selected measured processes, usually modestly. Evidence for broad, durable improvement in patient health is much less consistent, and effects depend on program design and context.

Does a higher performance score prove that care improved?

No. The increase may include better care, better documentation, changed coding, or denominator changes. Patient outcomes and audit data help show what moved.

Should incentives be removed if evidence is uncertain?

That decision requires more than an average review result. Policymakers should identify which measures produce value, watch for deterioration or harm during any change, and preserve useful infrastructure such as registries and recall systems. Withdrawal itself should be evaluated. Pay for performance is best understood as a behavior-and-measurement policy, not a guarantee of quality. Its success depends on what is rewarded, what is ignored, who bears the burden, and whether better scores translate into better lives.