Dr. Manouchehr Hessabi
← All writing
8 min readresearch methods · reproducibility · evidence

Why replication matters: what the 2026 projects measured

Why replication matters in science: what three 2026 Nature projects found, how reproducing a result differs from replicating it, and how to read one study.

By Manouchehr Hessabi, MD, MPH

A published finding is a claim that a pattern exists in the world. It is produced by one team, working with one sample, after one set of analytic choices. Replication, the attempt to find the same pattern again in new data, is how a field learns whether that claim describes the world or only the sample it came from.

On 1 April 2026, Nature published three coordinated meta-research projects that measured this at unusual scale. Taken together they support a conclusion that is undramatic and worth stating plainly: roughly half of the claims tested held up, and the ones that held up were substantially weaker than first reported. That is not evidence of widespread misconduct. It is closer to the expected behaviour of first estimates, and it changes how a careful reader should weigh any individual result.

This is educational and not a substitute for personal medical advice.

What is the difference between reproducing and replicating a result?

These two words are often used interchangeably and should not be.

Reproducibility asks a bookkeeping question: given the authors' own data and their own analysis code, does re-running the analysis produce the numbers they published? A failure here is a problem of transparency, documentation, or arithmetic. Nothing new is collected.

Replicability asks an empirical question: if new data are gathered under the same design, does the same pattern appear? A failure here says something about whether the finding generalizes beyond the original participants, setting, and moment.

The distinction matters because the two have different cures. Reproducibility responds to policy. Replication does not.

The clearest illustration comes from the third of the 2026 projects. Brodeur and colleagues reproduced the analyses of 110 articles published in leading economics and political science journals that require authors to share data and code, and found that more than 85% of published claims were computationally reproducible. In robustness checks, 72% of statistically significant estimates remained significant and in the same direction, and the median reproduced effect size was 99% of the published one.

That is close to what a well-functioning system should look like, and it was achieved largely by requiring the materials be available in the first place.

What did the 2026 Nature projects find?

The replication project, led by Tyner and colleagues, attempted replications of 274 claims of positive results from 164 quantitative papers published between 2009 and 2018 across 54 social and behavioural science journals. The replications were well resourced by the standards of this literature: they used original materials where available, followed protocols peer reviewed in advance, and had a median statistical power of 99.6%, meaning that if the original effect were real and the size originally reported, these studies would almost certainly have detected it.

Replications produced statistically significant results in the original direction for 151 of the 274 claims, or 55.1% (95% confidence interval 49.2 to 60.9). Weighted by paper, the figure was 49.3%. Rates varied only modestly by discipline, from 42.5% to 63.1%.

The more instructive result is about magnitude rather than pass rates. The median effect size, expressed as a Pearson correlation, was 0.25 in the original studies and 0.10 in the replications, an 82.4% reduction in shared variance. A claim can replicate in the technical sense and still describe a relationship far weaker than the one readers took away from the original paper.

The companion reproducibility audit, by Miske and colleagues, drew a stratified random sample of 600 papers published from 2009 to 2018 in 62 journals. The first finding arrives before any analysis: the authors of only 144 papers, 24.0%, made data available to assess. Among the datasets that could be assessed, 53.6% of papers were rated precisely reproducible and 73.5% at least approximately reproducible, meaning within 15% of the original effects or within 0.05 of the original P values. Reproducibility was higher for political science and economics, for more recent papers, and for papers in journals that require data sharing.

One honest caveat belongs with these numbers. The replication team applied thirteen different criteria for what counts as a successful replication, and those criteria produced rates ranging from 28.6% to 74.8%. "Replicated" is not a single well-defined threshold, and a careful reader of any replication study should ask which definition was used before comparing one project's headline to another's.

Is this a social science problem only?

No. The same shape appears wherever it has been looked for carefully.

The Open Science Collaboration replicated 100 experimental and correlational studies from three psychology journals in 2015. Ninety-seven percent of the original studies had reported statistically significant results; 36% of the replications did. Replication effects were about half the magnitude of the originals, and 47% of original effect sizes fell within the 95% confidence interval of the replication estimate.

In preclinical biomedicine, the Reproducibility Project: Cancer Biology repeated 50 experiments from 23 high-impact papers, assessing 158 effects. For positive effects, the median replication effect size was 85% smaller than the median original effect, and 92% of replication effect sizes were smaller than their originals. Combining positive and null effects, the overall success rate was 46%.

Why does the same pattern keep recurring across very different fields? Several mechanisms are associated with it, and none require anyone to behave badly. Journals have historically preferred to publish statistically significant results, a tendency known as publication bias. Estimates that clear a significance threshold in a small sample are selected from the noisier, larger end of the distribution of possible results. Analytic flexibility gives researchers many defensible paths through the same data, and the path that produces a clean result is the one most likely to reach print. Each of these inflates first estimates without any individual study being wrong on its own terms.

The cancer biology authors framed the correct interpretive posture better than most. A successful replication does not definitively confirm an original finding or its theoretical interpretation, and equally, a failure to replicate does not disconfirm a finding, but it does suggest that additional investigation is needed to establish its reliability. Non-replication is a signal to look harder, not a verdict.

Why does this matter for environmental and epidemiological research?

Most observational research cannot be replicated in the way an experiment can. Nobody re-randomizes a birth cohort. But the underlying logic transfers directly.

An association reported between an environmental exposure and a child health outcome in one cohort is best read as a hypothesis about other cohorts. It becomes durable knowledge when it is observed again with different participants, different exposure measurement, and a different structure of confounding. This is why consistency across studies sits among the viewpoints Austin Bradford Hill proposed for weighing whether an association is causal, discussed in more detail in the explainer on the Bradford Hill criteria, and why confounding is treated as a standing threat rather than a solved problem.

In observational work, replication takes recognizable forms: independent cohorts examining the same question, pooled and coordinated analyses across studies, analysis plans registered before the data are examined, and shared data that let others check the arithmetic. The NIH describes scientific rigor as "the strict application of the scientific method to ensure unbiased and well-controlled experimental design, methodology, analysis, interpretation and reporting of results," and its rigor and reproducibility guidance asks applicants to describe the experimental details that reviewers had too often been left to assume. That is the funding-policy expression of the same idea.

The limits deserve equal billing. Replication is expensive, and the resources spent repeating one study are not spent asking a new question. Some studies cannot be repeated at all, because the exposure, the population, or the moment no longer exists. And a failed replication can reflect a flawed replication as easily as a flawed original: a different population, a subtly altered protocol, or insufficient power will all produce a null. The 2026 projects are credible precisely because they addressed this, with high power, original materials, and protocols reviewed in advance. A reader should look for those same features before treating any single non-replication as decisive.

What should a reader do with a single study?

Three habits cover most of it.

  • Ask whether anyone else has found it. A result that has been seen by independent teams, with different methods, is in a different evidentiary category from one that has been published once.
  • Ask whether the data and code are available. In the 2026 audit, only about a quarter of papers made data available to check, and journals that required sharing did measurably better.
  • Expect the true effect to be smaller than the first estimate. This was the most consistent finding across all five projects described here, spanning psychology, economics, political science, and preclinical cancer biology.

None of this warrants cynicism about published research. A field that runs large, expensive, well-powered projects to audit itself, and then publishes results showing that roughly half its tested claims survive, is a field doing what science is supposed to do. The appropriate response for a reader is not distrust but patience: a claim is worth more once it has been seen twice.

Readers interested in the methodological work behind these questions can find the relevant peer-reviewed publications collected on this site.

About the author. Dr. Manouchehr Hessabi is a physician-epidemiologist and Senior Research Scientist at the BERD core of UTHealth Houston's Center for Clinical and Translational Sciences. See his peer-reviewed publications or research programs.