Benefit outcomes are usually named in the trial objective, powered in the sample-size calculation, and analyzed in a prominent table. Harms may appear in a short paragraph stating that treatment was “well tolerated.” That phrase is not an endpoint. A credible safety assessment explains which events were sought, when and how they were collected, and who was observed. It explains how events were coded and what uncertainty remains.
Begin with terminology#
An adverse event is an unfavorable medical occurrence after a participant receives an intervention; it need not have been caused by the intervention. An adverse reaction implies a reasonable causal relationship under the applicable framework. Authors sometimes use “side effect” loosely, so you need to find the formal definitions.
Seriousness is regulatory and consequence-based: death, life threat, or hospitalization or prolongation. It also covers substantial disability, congenital anomaly, or another medically important event under specified criteria. Severity is intensity, such as mild, moderate, or severe. A severe headache may not be serious; a mild symptom preceding hospitalization can be serious.
“Treatment-emergent” usually means onset or worsening after treatment begins within a defined window. That window, baseline comparison, and handling of recurring chronic symptoms should be stated.
The method of asking changes the answer#
Participants report more symptoms when given a checklist than when asked, “Any problems?” Scheduled laboratory tests detect asymptomatic abnormalities that spontaneous reporting misses. Clinician judgment, diary prompts, digital monitoring, and chart review each have different sensitivity.
If one group receives more visits or laboratory testing, surveillance itself can create more detected events. Open-label knowledge can influence participant reporting, clinician questioning, diagnostic workup, and attribution. Blinded outcome assessment helps but cannot compensate for unequal data collection.
CONSORT Harms 2022 asks trials to describe whether harms were prespecified or emerging. It asks for collection mode, timing, and ascertainment frequency. It asks for attribution, surveillance, and stopping rules. Without it, you cannot compare event percentages across two trials.
Prespecified harms deserve definitions#
Expected mechanism-based or clinically important harms should be named in the protocol with event definition, time window, and measurement instrument. The protocol should also give analysis and threshold. A composite such as “cardiovascular events” should list components and adjudication rules.
Emerging harms also matter. A rigid table limited to common expected events can hide an unexpected pattern. Trials need systems for coding and reviewing unsolicited events while preserving transparency about post hoc signals.
Laboratory “toxicity” requires a baseline, upper or lower limit, grade, persistence, and clinical context. Reporting only mean laboratory change can conceal a few dangerous extremes. Conversely, counting every transient mild abnormality equally can exaggerate concern.
Find the denominator#
Safety analyses often include everyone receiving at least one dose rather than every randomized participant; this can be appropriate for drug-attributable events, but post-randomization exclusion breaks the original group comparison if people fail to start treatment for prognostic reasons.
The report should show randomized, treated, assessed, and included counts by group. A percentage without a denominator is incomplete. If the denominator changes by event because tests are missing, state each one.
An “as-treated” safety analysis can misclassify events after switching and undermine randomization. Intention-to-treat, on-treatment, and as-treated views answer different questions. The protocol should define risk windows for events after discontinuation, crossover, dose interruption, and rescue treatment.
Participants and events are different units#
“Ten participants had at least one infection” measures risk of being affected. “There were 28 infections” includes recurrence. Both can matter. Reporting only participants discards recurrent burden; reporting only events lets a few participants dominate and violates independence if analyzed naively.
For recurrent events, methods such as rates or time-to-first event may be appropriate depending on the question and competing risks. So may mean cumulative function or recurrent-event models. Show how many people experienced multiple episodes.
Time-to-first analysis ignores later events. A person with five severe episodes and a person with one mild episode each count once; it may still be useful for a first-event estimand, but you need the burden data alongside it.
Time at risk and competing events matter#
Incidence proportion, participants with an event divided by participants observed, assumes comparable follow-up. If the intervention group stops treatment earlier, it has less time to accumulate on-treatment events. A lower proportion may reflect shorter observation rather than greater safety.
Event rates per person-time account for follow-up duration but make assumptions and can be dominated by recurrent events. Kaplan-Meier complements are often used for time-to-event risk, but death or another event that prevents the harm is a competing event; simply censoring it can overestimate cumulative incidence.
Treatment can also reduce survival time and therefore reduce opportunity to observe a nonfatal adverse event. That apparent reduction is not a safety benefit. Show deaths, discontinuations, and follow-up alongside event estimates.
Withdrawals are safety outcomes#
Report discontinuations due to adverse events, dose reductions, interruptions, and rescue treatments by reason and group. All-cause dropout blends efficacy, toxicity, preference, logistics, and administrative causes; it is not a pure tolerability measure.
After a participant stops assigned treatment, harms may continue. Restricting follow-up to active dosing misses withdrawal, rebound, and delayed organ toxicity. It misses pregnancy outcomes and events caused by long biological persistence.
Missing harm data can be informative. Someone hospitalized elsewhere or lost after a severe symptom may differ from a participant who completed every checklist. Sensitivity analyses should address plausible missing-event patterns.
Coding can hide or reveal patterns#
Trials often code investigator terms into a dictionary such as MedDRA, grouping them into preferred terms and system organ classes. Coding supports consistency but involves judgment. “Chest discomfort,” “chest pain,” and “myocardial ischemia” may be distributed across terms and levels.
Pooling related terms after data are known can either identify a syndrome or manufacture one, and reports should prespecify grouped queries where possible and disclose coding dictionary version, coding process, and any post hoc grouping.
A table that shows only events above a frequency threshold can hide a serious imbalance in a rare one, and nothing on the page tells you it is missing; CONSORT Harms recommends reporting important events regardless of frequency and explaining omissions, with access to fuller data.
Causality attribution can bias comparison#
Investigators may rate events unrelated, possibly related, or probably related. In an open-label trial, expectations can influence attribution. A table limited to “treatment-related” events can hide an objective imbalance because the subjective causality filter removed events unequally.
Present all adverse events and serious events, then attribution as an additional view. Randomized between-group differences can provide causal evidence even when individual-event attribution is uncertain.
A lack of statistical significance does not establish equal safety. Most trials have little power for specific harms, and many event comparisons raise multiplicity. Estimates and confidence intervals will tell you more than a column of unadjusted p-values.
Zero events does not mean zero risk#
If no event occurs among a small group, the upper confidence bound can still allow a clinically important risk, and a common rough rule says that with zero events in n participants, the upper 95% bound is about 3/n under simple assumptions. Zero among 100 does not exclude a risk near 3%.
Rare harms may emerge only after thousands or millions of treated people. Long latency, cumulative dose, and interactions are often underrepresented in preapproval trials. So are pregnancy, age, kidney or liver impairment, and genetic susceptibility.
Regulatory reviews, observational comparative studies, registries, spontaneous-reporting systems, and pharmacovigilance can extend the evidence. Each has strengths and biases. A signal is not automatically causal, and absence of a signal in passive reports is not proof of safety.
Benefits and harms may use different time horizons#
A six-week symptom benefit and a five-year malignancy risk cannot be placed in one short trial. Likewise, an early procedural complication may be balanced against a durable benefit. The follow-up you need depends on the mechanism and on the decision.
Harms can be patient-reported, functional, economic, and burdensome even when not medically serious. Sexual dysfunction, cognitive symptoms, and fatigue can determine whether a treatment is acceptable. So can treatment visits, monitoring, and withdrawal. Trial endpoints should reflect what participants consider important.
A practical appraisal sequence#
Locate the protocol and statistical plan. List prespecified and emerging harms, definitions, and collection methods. List schedule, coding, adjudication, and risk windows. Reconcile randomized, treated, and assessed denominators.
For each important harm, record participants affected, total events, and severity. Record seriousness, discontinuations, and follow-up time. Record effect estimate and confidence interval. Check recurrence, competing events, missingness, and post-treatment follow-up. Compare ascertainment across groups.
Then ask what the trial could not detect because of sample size, duration, exclusions, and selective tables. Integrate external safety evidence and the magnitude and durability of benefit. “Well tolerated” should be a conclusion earned by data, never a substitute for them.
Sources and further reading
- Junqueira and colleagues, CONSORT Harms 2022 Statement, Explanation, and Elaboration, BMJ (2023)
- EQUATOR Network, CONSORT Harms 2022 Reporting Guideline
- Hopewell and colleagues, CONSORT 2025 Statement, BMJ (2025)
- Ioannidis and colleagues, Better Reporting of Harms in Randomized Trials, Annals of Internal Medicine (2004)
- U.S. Food and Drug Administration, Safety Assessment for IND Safety Reporting Guidance
Questions and answers
Is every adverse event caused by treatment?
No. An adverse event occurs after treatment but may reflect underlying illness, chance, or another cause. Randomized differences and other evidence help assess causality.
Is severe the same as serious?
No. Severe describes intensity; serious is based on outcomes such as hospitalization, life threat, disability, or death under regulatory criteria.
Why not test every adverse-event p-value?
Trials are usually underpowered for individual harms, and testing many events creates false positives. Effect estimates, intervals, patterns, and external evidence are more useful.
Can equal event percentages prove equal safety?
No. Follow-up, discontinuation, surveillance, missingness, recurrence, and competing mortality may differ, and confidence intervals may be wide.
Where can rare harms be found?
They may require pooled trials, regulatory data, large observational studies, registries, and postmarketing surveillance. No single source is sufficient for every harm.