Calling a randomized trial “negative” should mean that it did not satisfy its prespecified rule for the primary outcome. It should not mean that the interventions were identical, that the estimate points nowhere, or that every secondary observation is false. ANDROMEDA-SHOCK is a useful example because its mortality estimate looked clinically important, yet the confidence interval crossed no difference and the p-value was just above the chosen threshold.
Reconstruct the exact comparison#
ANDROMEDA-SHOCK enrolled 424 adults with septic shock in 28 intensive care units across five countries. Participants were randomized equally to one of two stepwise resuscitation strategies during an eight-hour intervention period.
The peripheral-perfusion group targeted normalization of capillary refill time, measured by applying standardized pressure and timing return of color, and the comparator targeted normalization of serum lactate or a decrease greater than 20% every two hours. Both groups received structured hemodynamic assessment and treatment steps; this was not “capillary refill versus no resuscitation,” nor “ignore lactate versus use lactate.”
That framing matters. The trial evaluated management strategies, not the diagnostic accuracy of one measurement; effects could arise from how often clinicians reassessed, how treatment steps were triggered, how much fluid or vasoactive therapy followed, and how quickly a target normalized.
Read the primary result in full#
By 28 days, 74 of 212 participants in the peripheral-perfusion group and 92 of 212 in the lactate group had died: 34.9% versus 43.4%. The reported hazard ratio was 0.75 with a 95% confidence interval from 0.55 to 1.02 and p=0.06, and the risk difference was -8.5 percentage points with a 95% confidence interval from -18.2 to +1.2.
The point estimate favored the capillary-refill strategy. If it represented the truth, the absolute difference would be important. But the interval also included little or no benefit. Under the prespecified conventional two-sided alpha level, the trial did not demonstrate superiority for 28-day mortality. The authors' primary conclusion was restrained on exactly that point: the strategy did not reduce 28-day mortality compared with lactate-targeted resuscitation. That is a statement about the evidence standard, not a declaration that the true hazard ratio equals 1.
Why p=0.06 is not “almost proof”#
A p-value is the probability, under a specified null model and analysis, of data at least as incompatible with the null as those observed. It is not the probability that the intervention works, the probability that the null is true, or the chance the result will replicate.
The difference between p=0.049 and p=0.06 is not a sharp biological boundary. Yet prespecified thresholds serve an important governance function: they limit discretion after results are known. Rebranding p=0.06 as positive because it was close would weaken that protection.
The better response is to report estimate and interval. The trial was compatible with a substantial mortality benefit and with a small difference. It reduced uncertainty without eliminating it.
Power and precision are not excuses added afterward#
A trial's sample-size calculation uses assumptions about control risk, effect size, loss to follow-up, analysis, and acceptable errors. If the observed control rate, event count, or true effect differs, precision may be less than hoped. That does not make a missed endpoint positive.
“Underpowered” is often used loosely. Any nonsignificant result can be called underpowered for some smaller effect. The useful question is which effects the confidence interval rules out and which ones you still have to take seriously. ANDROMEDA-SHOCK did not establish no worthwhile effect; it also did not provide the prespecified confirmation of benefit. Stopping, changing the endpoint, or recalculating the threshold after seeing results would create more serious interpretive problems, and trial registration and a prespecified analysis plan help you separate intended inference from post-result explanation.
Secondary findings remain secondary#
The peripheral-perfusion group had a mean Sequential Organ Failure Assessment score at 72 hours one point lower than the lactate group, with a 95% confidence interval just below zero and p=0.045. Six other reported secondary outcomes did not differ significantly.
Secondary outcomes can illuminate mechanism and generate hypotheses, though they do not usually rescue a failed primary endpoint, especially when several are tested without a multiplicity plan that controls the chance of false-positive conclusions. Organ-dysfunction scoring also differs from mortality in clinical meaning and susceptibility to treatment and measurement choices. The safety report found no confirmed protocol-related serious adverse reactions. Absence of detected serious reactions in 424 participants over the trial period does not prove that either strategy is risk-free.
Lactate is important but biologically nonspecific#
Elevated lactate in septic shock is associated with worse outcomes and can reflect inadequate tissue perfusion. It can also rise through adrenergic stimulation and accelerated glycolysis, impaired hepatic clearance, mitochondrial and metabolic changes, medicines, seizures, or other processes. A persistent concentration is therefore not a direct readout of a single treatable deficit.
Treating a biomarker as a target creates a risk of “chasing the number.” If lactate remains elevated for a reason not corrected by more fluid, additional fluid can cause edema and organ harm. Conversely, a normal capillary refill time assesses a particular peripheral vascular response and does not guarantee adequate perfusion in every organ.
The measurements provide complementary information. Current sepsis guidance suggests using lactate trends in patients with elevated lactate while considering clinical context and using capillary refill time as an adjunct among other perfusion measures. A trial comparing protocols should not be simplified into “lactate is useless.”
Measurement quality matters for both targets#
Lactate requires specimen collection, processing, and laboratory or point-of-care analysis, and tourniquet time, delay, sampling site, and device can all matter. Trend interpretation depends on consistent timing relative to interventions.
Capillary refill time is immediate and inexpensive but can vary with pressure, duration, site, lighting, skin temperature, edema, pigmentation, ambient conditions, and observer technique. ANDROMEDA-SHOCK standardized the method and trained staff. Performance in routine practice may be more variable. An intervention that depends on repeated bedside measurement includes the measurement protocol as part of the intervention, so moving it to your own unit means reproducing that protocol, not merely writing “check capillary refill.”
Later Bayesian analyses do not replace the primary analysis#
Reanalyses can ask different questions, such as the probability of benefit under stated priors or patterns across subgroups. Bayesian approaches may give you uncertainty as a clinically intuitive probability, but their answers depend on the model and the prior. They do not retroactively change the prespecified frequentist success rule.
Subgroup signals are similarly fragile. Illness severity, baseline lactate, capillary refill, and treatment responsiveness may plausibly modify effect, but interaction estimates need adequate power, prespecification, and replication. Separate significance in one subgroup and nonsignificance in another does not prove a difference.
What ANDROMEDA-SHOCK-2 adds, and what it does not#
Published in 2025, ANDROMEDA-SHOCK-2 enrolled patients within four hours of septic shock at 86 centers in 19 countries; it compared a personalized hemodynamic resuscitation protocol targeting capillary refill time with usual care. The protocol used pulse pressure, diastolic arterial pressure, fluid-responsiveness assessment, and bedside echocardiography to tailor fluids, vasopressors, and inotropes.
Its primary endpoint was a hierarchical composite at 28 days: mortality first, then duration of vital support, then hospital length of stay. In 1,467 participants included in the primary analysis, the win ratio was 1.16, 95% confidence interval 1.02 to 1.33, p=0.04. The reported advantage was driven mainly by duration of vital support, not a demonstrated mortality difference.
That is important evidence for a broader personalized strategy. It is not a replication of capillary refill targeting versus lactate targeting, and it used a different comparator, intervention bundle, sample, and endpoint. It cannot turn the 2019 primary mortality result into a positive one. The two trials can be coherent: the first left uncertainty, while the second asked and answered a related but distinct question.
A disciplined reading of a negative trial#
Start with registration, protocol, and statistical plan. Write down the population, intervention, comparator, co-interventions, time horizon, and primary estimand before you look at the result. Verify the event numbers, effect measure, interval, threshold, missing data, adherence, crossover, and stopping rules.
Then separate three statements:
- What was demonstrated: whether the prespecified primary criterion was met.
- What remains plausible: the range of effects compatible with the interval and assumptions.
- What deserves study: secondary, mechanistic, and subgroup findings that require confirmation.
Avoid “no difference,” “trend toward significance,” and “would have worked with more participants” as substitutes for those statements. Finally, integrate later evidence without changing history. Science accumulates by answering new questions, not by editing old endpoints.
Sources and further reading
- Hernández and colleagues, ANDROMEDA-SHOCK Randomized Clinical Trial, JAMA (2019)
- ClinicalTrials.gov, ANDROMEDA-SHOCK Trial Record NCT03078712
- Evans and colleagues, Surviving Sepsis Campaign International Guidelines (2021), Intensive Care Medicine
- Hernández and colleagues, ANDROMEDA-SHOCK-2 Randomized Clinical Trial, JAMA (2025)
- ClinicalTrials.gov, ANDROMEDA-SHOCK-2 Trial Record NCT05057611
Questions and answers
Did ANDROMEDA-SHOCK show that capillary-refill targeting saves lives?
No. Its point estimate favored that strategy, but the prespecified 28-day mortality endpoint did not reach statistical significance and the confidence interval included little difference.
Did the trial show that lactate should be ignored?
No. It compared two resuscitation protocols. Lactate remains a prognostic and monitoring measure that must be interpreted in clinical context.
Is p=0.06 the same as no effect?
No. It means the result did not meet the selected significance rule. The estimate and confidence interval describe the effect sizes still compatible with the data.
Can a secondary endpoint make the trial positive?
Usually not. Secondary findings can be informative, but multiplicity and different clinical meaning prevent them from replacing a failed prespecified primary endpoint unless the design explicitly controlled for that pathway.
Did ANDROMEDA-SHOCK-2 confirm the first trial?
It supported a personalized capillary-refill-targeted bundle for a hierarchical composite versus usual care. Because the comparator, intervention, and primary outcome differed, it is related evidence rather than a direct replication.