A trial's data-sharing statement tells you whether anyone outside the study team can ever look at the raw participant data, but it is a statement of intent, not a guarantee of access. Since July 2018, journals that follow the International Committee of Medical Journal Editors (ICMJE) have required this short paragraph in every clinical trial report. It is one of the most useful and most misread lines in a paper, because a confident promise to share and an actual, downloadable dataset are two very different things.
Key points#
- Individual participant data (IPD) is the row-by-row record behind a published result: one line per patient, with baseline traits, treatment assignment, doses, outcomes, and dropouts.
- ICMJE journals require a data-sharing statement, but the rule mandates transparency about intent, not sharing itself. "We will not share" is a compliant answer if stated plainly.
- Most trial data moves through managed access: an independent repository, a written proposal, and a data-use agreement, not an open download.
- When outside teams do obtain the data, reanalysis sometimes reproduces the original conclusion and sometimes reverses it.
The layer beneath the published numbers#
Think of a trial report as the top floor of a building. It shows the finished rooms: means, event counts, hazard ratios, p-values, and the tables and figures assembled from them. IPD is the foundation underneath, the de-identified participant-level dataset where each row is a single enrolled person and each column is something measured about them.
A dataset on its own is not enough to check a trial. Serious sharing bundles the documents that make the numbers interpretable: the study protocol, the statistical analysis plan, and, increasingly, the analytic code and a blank consent form. The reason is practical. Without the protocol and the analysis plan, a reviewer cannot tell which comparisons were planned before the results arrived and which were chosen afterward, and that distinction often decides whether a finding is trustworthy. A spreadsheet without its instructions is a set of numbers no one can fully audit.
What the statement actually commits to#
The ICMJE finalized this rule in 2017 with two staggered dates. Manuscripts submitted on or after July 1, 2018 must carry a data-sharing statement, and trials that began enrolling on or after January 1, 2019 must file a data-sharing plan at the moment they register.
The statement has to answer a fixed checklist: whether de-identified IPD will be shared, which specific data, which supporting documents travel with it, when the data becomes available and for how long, and through what mechanism and eligibility criteria, meaning who may request it and for what purpose. The committee is direct that leaving the question open is not allowed; authors have to commit to a position.
One point is easy to overlook. An earlier 2016 proposal would have made sharing de-identified IPD a condition of getting published at all. The final policy pulled back from that. As the committee explained in its editorial in the New England Journal of Medicine, what the rule requires is honesty about intent, not a duty to share. A trial team can fully satisfy the rule by clearly stating that it will not release its data. Compliance means the box is checked, not that the data is available.
Who is actually allowed to look#
"Sharing" spans a wide spectrum, and the two ends are worlds apart.
At the open end, a de-identified dataset sits in a public repository and anyone can download it. This is uncommon for clinical trials because of the privacy risk in patient-level records.
Far more typical is managed access. Here an independent platform holds the data, and an outside researcher has to submit a written proposal, sign a data-use agreement, and clear a review before receiving a gated copy that stays inside the platform. Repositories such as Vivli and the Yale YODA Project run this way, sitting as neutral brokers between the original sponsor and the person asking to reanalyze. This is the "mechanism and criteria" the ICMJE statement is meant to describe. It is a controlled, logged process built to protect participant privacy, not an open door.
That design has a useful side effect. The people best placed to stress-test a trial are often independent groups with no stake in the original conclusion, and managed access is precisely what lets them in without ever exposing a patient's identity.
Why raw data can confirm or overturn a result#
Reanalysis is the whole point. When independent investigators rebuild a trial from its IPD, they sometimes land on a different answer than the published one, and the reason is instructive.
A 2014 study in JAMA by Ebrahim and colleagues looked at 37 published reanalyses of randomized trials. About a third of them, 13, reached a conclusion different from the original about which patients should be treated. The underlying data never changed. The analytic choices did, and those choices moved the bottom line. That is not evidence of fraud; it is evidence that how you analyze a dataset can be as consequential as what the dataset contains.
Sharing can also work in the opposite direction, by shoring a result up. A 2018 study in The BMJ by Naudet and colleagues examined trials in The BMJ and PLOS Medicine, two journals with strong sharing policies. Every trial had a data-sharing statement, yet the authors could actually obtain usable data for only 17 of 37, roughly 46 percent. Where the data did come through, most reanalyses reproduced the original findings. Two lessons live inside that single study: a statement is not the same as accessible data, and when data is genuinely available, published conclusions often hold up.
Reading a data-sharing statement well#
For anyone weighing a trial, this paragraph rewards a careful read without inviting overinterpretation.
A firm "yes, de-identified IPD will be available through a named repository, alongside the protocol and analysis plan" signals a team that expects to be checked and has prepared for it. A vague or negative statement is not proof that something is wrong, since real constraints exist around privacy, consent, and small samples, but it does mean the published analysis is the only version anyone will ever see. And a pledge to share is a starting line, not proof that a dataset ever changed hands. Open trial data delivers its value only when someone requests it, receives it, and looks.
Sources and further reading
Questions and answers
Does a data-sharing statement mean I can download the data?
Usually not. Most trials use managed access, where an independent repository releases a gated copy only after you submit a proposal and sign a data-use agreement. Open, public datasets are the exception for clinical trials, mainly because patient-level records carry a privacy risk.
Can a trial refuse to share and still comply with the rule?
Yes. The ICMJE requires a clear statement of intent, not sharing itself. A trial that plainly declares it will not release its data meets the rule. That is why the statement tells you about transparency, and only sometimes about access.
Why does reanalysis of the same data sometimes reach a different conclusion?
Because analytic decisions, such as which patients to include, which outcome to prioritize, and how to handle missing data, can shift the result even when the numbers are identical. This is one reason the protocol and analysis plan matter as much as the dataset itself.