Research Not Reproducible? A Practical Guide
Research not reproducible occurs when an independent group, given the same data, code, and described methods, cannot obtain the same result, a problem documented across disciplines that computational scientists can attack directly because software, not wet-lab technique, is usually the failure point. Standards such as the 2019 National Academies report Reproducibility and Replicability in Science define the working vocabulary, and the practical fixes are concrete: version control, pinned dependencies, containerization, and build automation.
- Reproducibility and replicability are distinct: reproducibility means re-running the same analysis on the same data and getting the same numbers; replicability means a new study reaches a consistent conclusion.
- Most computational research not reproducible traces to four causes: missing or mutated code, unpinned software environments, undocumented data provenance, and hidden manual steps.
- A reproducible build is a deterministic transformation from source inputs to outputs — the same idea as reproducible builds in software packaging, applied to analysis pipelines.
- The minimum viable fix is a repository with a lockfile, a single entry point (
make, Snakemake, or a script), and a README that a stranger can follow without asking you questions. - Containers (Docker, Apptainer/Singularity, Conda environments) solve environment drift; they do not solve undocumented data cleaning or nondeterministic algorithms.
- Reproducibility is a spectrum, not a binary — aim for “a competent stranger can regenerate your figures,” not perfection.
What “Reproducible” Actually Means (and What It Doesn’t)
Reproducible research is the practice of packaging a study so that an independent researcher can regenerate its results from the original data and code. The word carries a specific technical meaning that differs from everyday usage: in the dictionary sense, “reproducible” simply means capable of being produced again, but in computational science it implies determinism, provenance, and enough documentation that the regeneration does not depend on tacit knowledge held by the original authors.
Reproducibility in research falls into a family of related terms that are frequently confused. The 2019 National Academies consensus study distinguishes computational reproducibility (same data, same code, same results) from replicability (new data, consistent findings).
Statisticians add repeatability, meaning the same analyst gets the same result twice. A study can be reproducible but not replicable – the code faithfully regenerates a result that a new experiment fails to confirm – and this distinction is important when diagnosing a failure.
The variant spelling “reproducable” often appears in search queries and informal writing, but Standard English uses “reproducible”. The suffix attaches to the root “reproduce” in the form “-ible”, corresponding to “producible” and “deducible”. Style guides and dictionaries only list “reproducible”; treat “reproducable” as a spelling error, even though you will encounter it in old forum posts and in non-native-English manuscripts.
Why Research Is Not Reproducible: The Four Failure Modes
Missing or mutated code accounts for a large portion of failed reproduction attempts. A common scenario: the published article describes an analysis, the corresponding author has left the lab, and the repository either does not exist, contains an outdated version, or references a private dataset. Code that “works on my machine” but has no version history cannot be audited, and without a commit hash, there is no way to know which version produced the published figure.
Related: — Project-based data-science paths with a guided terminal and real datasets.
Unpinned software environments stop reproduction silently. A NumPy pipeline written with version 1.20 may produce different floating point results, deprecation warnings, or outright errors under 2.x.
Molecular dynamics trajectories generated with one GROMACS or LAMMPS build may differ from another when compiler flags, FFT libraries, or parallel decomposition change. The fix is to save the exact versions — a requirements.txt with hashes, a conda env export, a renv.lock for R, or a container image digest — not just a list of package names.
Undocumented data provenance is the third failure mode. Raw instrument output, simulation restart files, and HDF5 archives often go through multiple cleanup and conversion steps that are never shown in the methods section. If a collaborator manually filtered outliers in a spreadsheet, this step is invisible and non-reproducible. Provenance means recording where each input came from, what transformations were applied, and in what order.
Worth a look: — One subscription for university-backed Python and data-science certificates.
Hidden manual steps are the fourth. Clicking through a GUI to export a plot, renaming files by hand, or copying numbers between tools all break automation. The test is simple: could someone who has never met you run your pipeline end to end without a phone call? If the answer is no, the manual steps are the problem.
How to Make Research Reproducible: A Practical Workflow
Version control is the foundation. Git (or Mercurial for existing projects) tracks every change made to code, manuscripts, and small configuration files. Commit early, write meaningful posts, and mark the exact commit that matches a submitted article. For large binary artifacts (trajectories, HDF5 files, trained model weights) use Git LFS, DVC or Zenodo/figshare repositories with DOIs rather than committing gigabytes to the repository.
Pinning dependencies comes next. Python projects must provide a lock file (pip-compile, Poetry or uv); R projects should use “renv”; Julia projects have “Project.toml” and “Manifest.toml”. Pin transitive dependencies, not just direct dependencies, because a deep change at three levels can change the digital output. Record the operating system, compiler, and any hardware-specific libraries (BLAS, CUDA) that affect the results.
A single entry point turns a folder of scripts into a pipeline. The classic tool is make: a Makefile declares targets, dependencies, and commands, and make rebuilds only what changed. This is the essence of reproducible research with Make — the build graph documents the workflow and enforces order. Modern alternatives include Snakemake, Nextflow, and targets (for R). Any of them beats a README that says “run script1.py, then script2.py, then…”.
Containerization captures the whole environment. Docker images, Apptainer/Singularity images (common on HPC clusters where Docker is unavailable), and Conda environments each solve environment drift. A container with a pinned digest is the strongest guarantee that your analysis runs identically on a laptop, a cluster, and a reviewer’s machine. The trade-off is image size and build time; a lean base image and multi-stage builds keep both manageable.
The documentation comes full circle. A README should indicate the software requirements, the exact command to reproduce each figure, the expected execution time and the expected result. A CITATION.cff file or a codemeta.json helps others cite software correctly. When a step is truly interactive, describe it specifically enough that a reader can repeat it.
Reproducible Research with R: A Concrete Example
R has one of the most mature reproducibility ecosystems, which is why “reproducible research in R” is a common search. The renv package snapshots package versions into a lockfile, so renv::restore() rebuilds the exact library a project was developed against. R Markdown and Quarto weave code, output, and prose into a single document that regenerates when knitted, eliminating copy-paste errors between analysis and manuscript.
A minimal R project structure looks like this: a renv.lock for dependencies, an R/ directory for functions, a data-raw/ directory for ingestion scripts, a data/ directory for processed objects, and a Makefile or _targets.R that orchestrates the build. The targets package extends make-style dependency tracking to R, skipping steps whose inputs have not changed. For more in-depth treatment, the book Reproducible Research with R and RStudio by Christopher Gandrud is a standard reference, and R Programming for Data Science by Roger Peng and the Johns Hopkins reproducibility course materials cover the workflow from start to finish.
The same principles carry over to Python. A pyproject.toml with pinned dependencies, a Makefile or Snakemake workflow, and a Quarto or Jupyter Book manuscript give you the equivalent of the R stack. The tooling differs; the discipline does not.
What Are Reproducible Builds?
Reproducible builds are a software-engineering practice in which compiling the same source code with the same toolchain yields a byte-for-byte identical binary. The concept originated in the free-software world — the Debian and Tor projects pioneered it — and is now a formal standard maintained by the Reproducible Builds project. The motivation is trust: if anyone can rebuild a binary and get the same hash, they can verify that the distributed binary matches the published source and contains no hidden modifications.
The analogy to scientific pipelines is direct. A reproducible analysis build takes source inputs (raw data, code, configuration) and produces outputs (figures, tables, statistics) deterministically.
Nondeterminism breaks this, making research not reproducible: random seeds not fixed, parallel reductions that sum in varying order, timestamps embedded in output files, and floating-point operations that depend on thread count all cause run-to-run variation. Fixing seeds, sorting before reduction, and stripping timestamps are standard remedies.
Not all sources of non-determinism can be eliminated. Molecular dynamics with GPU acceleration can produce slightly different trajectories on different hardware due to different floating point rounding, and Monte Carlo methods are inherently stochastic. The honest approach is to document the expected variation, report it, and ensure that the scientific conclusion is robust to it – not to claim a bit-level identity that you cannot provide.
Criteria for Judging Whether a Study Is Reproducible
| Criterion | Weak | Strong |
|---|---|---|
| Code availability | ”Available on request” | Public repository with a tagged release and DOI |
| Environment | Package names only | Lockfile or container with pinned digest |
| Entry point | Numbered scripts, manual order | make, Snakemake, or targets workflow |
| Data provenance | Methods paragraph | Documented ingestion scripts and checksums |
| Determinism | Unseeded randomness | Fixed seeds, documented expected variation |
| Documentation | Assumes author knowledge | Stranger can regenerate every figure |
Use this as a self-audit before submission to ensure your research is not reproducible. Most journals and funders now require a data and code availability statement, and some — including several AGU, PLOS, and Nature portfolio journals — require code to be deposited in a repository with a DOI. Meeting the “strong” column is increasingly a condition of publication, not a bonus.
Common Objections and Honest Trade-offs
The cost of time is the most cited objection, and it is real. Properly packaging a pipeline can take days for an already “finished” project. The counterargument is that the cost of not doing it is paid later, often by a student or collaborator spending weeks reverse-engineering an analysis. A pragmatic middle path: invest immediately in version control and a lockfile, and only containerize projects that you expect others to reuse.
Sensitive or restricted data makes sharing difficult. Clinical, proprietary and certain instrument data cannot be published publicly. In these cases, reproducibility means publishing the code, the schema, and a synthetic or de-identified sample, along with a clear description of the access procedure. The goal is that a qualified researcher with legitimate access can regenerate the results.
Legacy code is another constraint. Fortran and C simulations from the 1990s may not compile on modern systems without patches. Wrapping them in a container with an old toolchain is often easier than rewriting them, and it preserves the original numerics. Document what patches you applied and why.
Finally, reproducibility is not the same as correctness. A pipeline can be perfectly reproducible and still be wrong — a bug faithfully reproduced is still a bug. Reproducibility is a prerequisite for scrutiny, not a substitute for it. Peer review, independent replication, and sensitivity analysis remain necessary to ensure research is not reproducible only in its errors.
Frequently Asked Questions
What does it mean when research is not reproducible?
Research is not reproducible when an independent researcher, using the same data and code, cannot regenerate the published results. The cause is usually missing code, an unpinned software environment, undocumented data processing, or manual steps that were never automated. This is a property of the artifacts, not a judgment on the honesty of the original authors.
What is the difference between reproducibility and replicability?
Reproducibility means re-running the same analysis on the same data and obtaining the same result. Replicability means conducting a new study – new data, new samples – and reaching a consistent conclusion. The 2019 National Academies report treats these concepts as separate concepts, and a study can be reproducible even if it fails to be replicated.
How do I make my research reproducible?
Start with version control (Git) and mark the commit that corresponds to your article. Pin dependencies with a lock file or container image. Provide a single entry point such as a “Makefile”, Snakemake or “targets” workflow. Document the data provenance and expected results in a README file. Test everything on a clean machine before submitting.
Is “reproducable” a correct spelling?
No. Standard English spells the word “reproducible,” formed from “reproduce” plus the suffix “-ible.” Dictionaries and style guides list only “reproducible.” The variant “reproducable” appears in informal writing and some non-native-English text but is not accepted in formal scientific writing.
What are reproducible builds in scientific computing?
Reproducible builds are builds that produce identical outputs from identical inputs, a standard formalized by the Reproducible Builds project for software packaging. In science, the same principle applies to analysis pipelines: fixed random seeds, deterministic reductions, and pinned toolchains allow results to be regenerated exactly. When hardware-dependent floating-point variation is unavoidable, document the expected range.
Does reproducibility require sharing all my data?
Not always. Restricted, clinical, or proprietary data can be handled by publishing code, schemas, and synthetic samples alongside a documented access procedure. Many funders and journals require a data availability statement rather than open data. The test is whether a qualified researcher with legitimate access could regenerate your results.
Further Reading and Tools
The National Academies’ Reproducibility and Replicability in Science (2019) is the authoritative conceptual reference for when research is not reproducible. The Reproducible Builds project documents the software-engineering standard.
For workflow tooling, see the documentation for GNU Make, Snakemake, Nextflow, targets, renv, and Quarto. The Turing Way is an open community handbook covering reproducible research practices across disciplines, and the Software Sustainability Institute publishes practical guidance for research software engineers.
P.S. A few readers have asked which interactive course platform we actually reach for — it's DataCamp; if you want the current details.
Frequently asked questions
What does it mean when research is not reproducible?
Research is not reproducible when an independent researcher, using the same data and code, cannot regenerate the published results. The cause is usually missing code, an unpinned software environment, undocumented data processing, or manual steps that were never automated. This is a property of the artifacts, not a judgment on the honesty of the original authors.
What is the difference between reproducibility and replicability?
Reproducibility means re-running the same analysis on the same data and obtaining the same result. Replicability means conducting a new study – new data, new samples – and reaching a consistent conclusion. The 2019 National Academies report treats these concepts as separate concepts, and a study can be reproducible even if it fails to be replicated.
How do I make my research reproducible?
Start with version control (Git) and mark the commit that corresponds to your article. Pin dependencies with a lock file or container image. Provide a single entry point such as a “Makefile”, Snakemake or “targets” workflow. Document the data provenance and expected results in a README file. Test everything on a clean machine before submitting.
Is 'reproducable' a correct spelling?
No. Standard English spells the word 'reproducible,' formed from 'reproduce' plus the suffix '-ible.' Dictionaries and style guides list only 'reproducible.' The variant 'reproducable' appears in informal writing and some non-native-English text but is not accepted in formal scientific writing.
What are reproducible builds in scientific computing?
Reproducible builds are builds that produce identical outputs from identical inputs, a standard formalized by the Reproducible Builds project for software packaging. In science, the same principle applies to analysis pipelines: fixed random seeds, deterministic reductions, and pinned toolchains allow results to be regenerated exactly. When hardware-dependent floating-point variation is unavoidable, document the expected range.
Does reproducibility require sharing all my data?
Not always. Restricted, clinical, or proprietary data can be handled by publishing code, schemas, and synthetic samples alongside a documented access procedure. Many funders and journals require a data availability statement rather than open data. The test is whether a qualified researcher with legitimate access could regenerate your results. Further Reading and Tools The National Academies' Reproducibility and Replicability in Science (2019) is the authoritative conceptual reference for when research is not reproducible. The Reproducible Builds project documents the software-engineering stand
Learn Python by coding in your browser
Interactive Python and data-science courses you code directly in the browser