A methods section tells you what comparison the investigators planned, who contributed data, how outcomes were measured, and which analysis was meant to carry the main claim, and if those pieces do not line up, a polished result cannot repair the mismatch.
Reading methods first also reduces hindsight bias. Once you know which result was favorable, an improvised analysis can seem inevitable; before seeing that result, it is easier to ask whether the same analysis would have looked fair if the finding had gone the other way.
Key points#
- Start with the research question and decide which design could answer it.
- Identify the exact population, intervention or factor, comparator, outcome, and time horizon.
- Find the prespecified primary outcome and analysis, then distinguish them from exploratory work.
- Check how assignment, masking, missing data, protocol deviations, and measurement could create bias.
- Treat transparent reporting as necessary for appraisal, not as proof that the study itself was well designed.
Reconstruct the question before judging the answer#
A useful first pass is to write the question in a PICO-like form: population, intervention or factor, and comparator. Then add outcome and time. In a randomized treatment trial, this might be adults with a defined condition, assigned to a drug or placebo, followed for 24 weeks, with change in a validated symptom score as the primary outcome. In an observational study, the intervention becomes a naturally occurring factor, and the method must address why people with that factor differ from those without it.
Then ask whether the design can support the wording of the conclusion. Randomization can estimate a causal effect when assignment is preserved and follow-up is adequate. A cohort can estimate associations over time but remains vulnerable to confounding. A cross-sectional survey can describe what coexists at one moment, not which feature caused the other, and a diagnostic-accuracy study can estimate how a test classifies people against a reference standard, but it does not by itself show that using the test improves health.
This claim-to-design check catches many problems quickly. If a paper makes a treatment claim from a before-and-after case series, the main weakness is not a missing statistical adjustment. It is the lack of a concurrent comparison capable of separating treatment effect from natural history, co-interventions, and regression to the mean.
Find out who was actually studied#
Eligibility criteria define the population to which the estimate most directly applies. Look beyond the condition name. Age limits, disease severity, organ function, previous treatments, pregnancy status, language requirements, and ability to attend visits can create a study population unlike the people who will later face the decision.
Recruitment method matters too. A consecutive clinic sample, a volunteer registry, an insurance database, and a social-media campaign select people differently. A flow diagram should account for those assessed, excluded, assigned, followed, and analyzed; large numbers do not remove selection bias if entry into the dataset depends on both the factor and the outcome.
For multisite studies, ask whether sites represent the settings in which the result will be used, because a model developed at tertiary referral centers may perform differently in primary care, where prevalence, testing patterns, and documentation differ. Generalizability is not a yes-or-no label. It is a judgment about how closely the study conditions match the target decision.
Audit allocation, masking, and the comparator#
In a randomized trial, sequence generation and allocation concealment solve different problems. A genuinely random sequence prevents systematic assignment. Concealment prevents recruiters from predicting the next assignment and steering particular participants toward it. The methods should explain both.
Masking can reduce changes in care, reporting, and outcome assessment caused by knowing the assignment. It is not always feasible, especially for procedures or behavioral interventions. When it is absent, ask which outcomes are most vulnerable. Death is less open to interpretation than a clinician-rated symptom scale, but even an objective outcome can be influenced by unequal co-interventions or follow-up.
The comparator must answer the practical question. Placebo control can estimate efficacy under controlled conditions. Active control may better answer which option to choose. Usual care must be described because it varies by place and time. A weak comparator can make a modest intervention look impressive without showing it is preferable to current practice.
Trace the outcome from definition to measurement#
The primary outcome should be named, timed, and operationalized. “Cardiovascular events” is not enough. Readers need the components, adjudication rules, and observation window. They need handling of recurrent events and whether adjudicators were masked. For a questionnaire, look for its scale, scoring, validation, and clinically meaningful interpretation.
Composite outcomes need special attention. A treatment may reduce a frequent, less serious component while leaving death or disabling events unchanged. Time-to-event outcomes require definitions of time zero, censoring, and competing events. Surrogate outcomes can be valuable, but an improved laboratory measure is not automatically an improved patient outcome.
Measurement should be comparable across groups. If one group is monitored more intensively, it may accumulate more diagnoses even when underlying disease is the same, and if an algorithm labels the outcome, the algorithm's validation and whether it used information related to treatment assignment belong in the appraisal.
Separate the planned analysis from the available analysis#
Registration, the protocol, and a dated statistical analysis plan help distinguish prespecified choices from choices made after results were visible. Compare the registered primary outcome, time point, sample size, subgroups, and model with the paper. Changes may be justified, but they should be dated and explained.
The sample-size calculation reveals the expected effect, event rate, and variability. It reveals power and anticipated loss to follow-up. It is not a certificate of adequacy. If assumptions were wrong, the achieved precision matters more than whether the target enrollment was reached.
For randomized trials, intention-to-treat analysis usually preserves the comparison created by randomization. Per-protocol analysis can address a different question about adherence, but adherence is often related to prognosis. ICH E9(R1) encourages authors to define the estimand: the population, treatment condition, and outcome variable. The estimand also names the summary measure and strategy for events such as stopping treatment or using rescue therapy. This makes clear what effect the analysis is trying to estimate.
Missing data deserve more than the phrase “multiple imputation was used.” Ask why values were missing, whether missingness differed by group, which variables informed the model, and whether sensitivity analyses considered departures from its assumptions. No statistical method can recover information without assumptions.
Appraisal limits and a compact reviewer checklist#
Before accepting the abstract's conclusion, ask:
- Does the design match the causal, prognostic, diagnostic, or descriptive claim?
- Are the target population and recruitment pathway clear?
- Is the comparator clinically meaningful and sufficiently described?
- Were assignment and outcome assessment protected from foreseeable bias?
- Is the primary outcome exact, patient-relevant, and measured consistently?
- Do registration, protocol, and report agree on the main analysis?
- Are missing data, deviations, multiple testing, and model assumptions handled transparently?
- Are effect sizes shown with uncertainty, not only p values?
- Do sensitivity analyses test the assumptions most capable of changing the answer?
- Does the conclusion stay within the population, outcome, and time period actually studied?
CONSORT and other reporting guidelines help authors disclose these elements. Cochrane's RoB 2 tool helps reviewers judge bias in randomized trials. Neither can turn an unsuitable design into a suitable one. Their value is to make the chain of reasoning inspectable.
Sources and further reading
Questions and answers
Should I always read the methods before the results?
It is often useful. You can decide what would count as a fair comparison before a favorable result influences that judgment. You will still move back and forth between sections during a full appraisal.
Does registration prove that every analysis was prespecified?
No. Registration is a dated record, but entries can be incomplete or changed. Compare the registry history with the protocol, statistical analysis plan, and final report.
Is a long methods section a sign of high quality?
Not by itself. Detail improves auditability, but quality depends on whether the methods answer the question while limiting bias. A clearly reported weak design remains weak.
What if an important method is missing from the paper?
Check supplements, the protocol, and the analysis plan. If the information remains unavailable, treat the corresponding risk of bias as uncertain rather than assuming the preferred method was used.