If you want to know how to make your analysis reproducible, here is the short version: a stranger should be able to take your project folder, run one command, and end up with the same numbers, tables and charts you reported. That requires three things most people skip, namely untouched raw data, every manual step turned into a script, and software versions pinned so nothing quietly changes under you.
The reason this matters is timing. Reproducibility only becomes urgent months later, when an editor asks you to move a number, a reviewer challenges a method, or a data correction forces you to rebuild a chart you already published. By then the reasoning is gone and only the code is left.
The good news is that this is production work, not academic virtue signalling. You do not need elegant code or perfect style. You need the steps preserved, the assumptions written down, and a test you can run to prove the pipeline still works.
Table of Contents
- What You Need
- Step-by-Step: How to Make Your Analysis Reproducible
- 1. Define the Question and Expected Result
- 2. Preserve the Input Data
- 3. Separate Raw and Cleaned Data
- 4. Write the Analysis as an Ordered Script
- 5. Record the Environment and Dependencies
- 6. Generate the Final Output and Record Parameters
- 7. Test the Workflow From a Clean State
- 8. Package the Analysis for Another Person
- Common Mistakes
- Frequently Asked Questions
- What does it mean for an analysis to be reproducible?
- Do I need a Jupyter notebook or a reproducible manuscript to make my analysis reproducible?
- How do I stay reproducible when the data is embargoed, proprietary or cannot be shared?
- How much documentation is enough for a reproducible analysis?
- Must a reproducible analysis produce identical numbers every time?
- How do I check whether an analysis is really reproducible before I rely on it?
- Conclusion
What You Need
Six materials, and nothing exotic:
- The source data, lawfully obtained — a downloaded file, an API response, a public register extract, or a copy of a file someone supplied you.
- A written research question, narrow enough that you can say what result would confirm or kill it.
- Scripts in a real language, or at minimum an ordered, named set of documented transformations in a spreadsheet or notebook.
- Pinned software versions, recorded in a dependency file rather than remembered.
- Documentation, mainly a README that a colleague could follow without emailing you.
- A predictable folder structure, so anyone can find the source data, the intermediate files and the final output.
If you are starting from an old folder of final_FINAL_v3 files, retrofitting is normal work. You do not have to rebuild from zero; you work backwards from the published number until you find the script that produced it.
Step-by-Step: How to Make Your Analysis Reproducible
1. Define the Question and Expected Result
Write the question before you touch the data, then state in one sentence what result would count as confirmation. This stops the analysis from wandering after an interesting outlier.
Instead of “look at rents”, write something like: “Among renter households in X county between 2022 and 2025, what share of listed rents rose by more than 15 percent year over year, using the quarterly register extract, excluding commercial properties.” Now define the population, the period, the units, the exclusions and the intended output — one table row per county.
How do you know it worked? Someone reading only your question can name the input file, the filters and the shape of the answer before seeing any code.
2. Preserve the Input Data
Store the input exactly as you received it and never edit it in place. Renaming a column or deleting a bad row in the downloaded file destroys your ability to prove where the number came from.
Record four things next to it: where it came from, when you pulled it, what you did to retrieve it, and any licence or usage constraint. For a newsroom project that means an API response saved verbatim with its request parameters and a timestamp, a scraped page saved as the HTML you actually parsed, and a supplied file stored untouched with the sender’s name and delivery date.
| Source type | Snapshot it as | Record alongside it |
|---|---|---|
| Public API | Raw JSON or CSV response | Request URL, parameters, pull timestamp |
| Scraped page | Page HTML plus the selector or parser used | URL, timestamp, user agent if it mattered |
| Downloaded file | CSV or Excel file as served | URL, timestamp, checksum |
| Emailed or handed over file | Untouched copy of the original | Sender, delivery date, licence terms |
How do you know it worked? Nothing under the raw folder has a modification date later than the day you downloaded it.
3. Separate Raw and Cleaned Data
Build one cleaning script that reads the preserved source and writes a new file into a cleaned folder. Two folders, two jobs, no editing across the boundary.
Inside that script, handle missing values explicitly rather than by deletion, drop duplicates on a declared key, normalise date formats to one standard, map category labels through a lookup table instead of by hand, and print row counts before and after so a surprise shows up in the log. Write assertions if your language has them.
A folder layout that survives contact with a second person:
rent-registry-2025/
data/
raw/ # downloaded files, never edited
interim/ # cleaned and merged, regenerable
output/ # final tables, charts, CSVs
scripts/
01_fetch.py
02_clean.py
03_analyse.py
04_chart.py
README.md
requirements.txt
Makefile
Add a rule to version control that keeps the code and the outputs while ignoring the bulky regenerable middle:
data/raw/*
!data/raw/.gitkeep
data/interim/
.ipynb_checkpoints/
__pycache__/
.venv/
.Rproj.user/
.Rhistory
.DS_Store
Keep the raw folder out of the repository when files are large or licensed, but keep the script that fetches it and a manifest describing what belongs there. Provenance you cannot commit is still provenance if the retrieval is scripted.
How do you know it worked? Delete the cleaned file, run the cleaning script, and get the same row count you recorded the first time.
4. Write the Analysis as an Ordered Script
Break the analysis into small scripts that run in a fixed order, each with one job. A script that fetches, cleans, joins and plots is four scripts wearing a trench coat, and you cannot rerun half of it after a source update.
Load the cleaned file explicitly at the top of each script, use relative paths so the project runs from any machine, and comment the decisions a reader would question rather than narrating every line.
# 03_analyse.py
import pandas as pd
df = pd.read_csv("data/interim/clean_rents.csv")
# Council areas that publish fewer than 200 listings a quarter
# are too noisy to report, so they are excluded from the denominator.
counts = df.groupby(["area", "quarter"]).size().reset_index(name="listings")
eligible = counts[counts["listings"] >= 200]["area"]
d = df[df["area"].isin(eligible)]
d["yoy_change"] = d["rent"] / d["rent_lag"] - 1
summary = (d[d["yoy_change"] > 0.15]
.groupby("area")
.agg(share_above_15pc=("yoy_change", lambda s: len(s) / len(d[d.index.isin(s.index)]))))
.reset_index())
summary.to_csv("data/output/share_above_15pc.csv", index=False)
If you work in notebooks, keep exploration in notebooks and keep the published pipeline in scripts. Notebooks carry hidden state, and running cells out of order then produces numbers that no fresh run can reproduce. Practitioners on data-science forums describe notebooks as fine for exploring and unacceptable as the pipeline of record.
How do you know it worked? Someone else can change one threshold, rerun a single script, and see only the affected numbers move.
5. Record the Environment and Dependencies
Most broken analyses break for boring reasons: a package updated, the interpreter moved on, and a default changed. Pin everything the result depends on.
# requirements.txt
pandas==2.2.3
numpy==2.1.2
geopandas==1.0.1
requests==2.32.3
python-dateutil==2.9.0.post0
# R
renv::init()
renv::snapshot()
# renv.lock is committed; renv/ is not in .gitignore
Modern Python tooling can go further and lock resolved versions, including transitive ones, which catches the dependency that arrives through another package. A conda environment file works the same way for teams that already use conda.
Set any random seed once, at the top of the script that samples, and pass it down. Run order and locale can change results too, so avoid relying on a path like C:Usersyou or on your system’s default date format.
How do you know it worked? On a fresh machine, the dependency file plus the code installs and runs without you answering a single question.
6. Generate the Final Output and Record Parameters
The final artefact — table, chart, map, statistic — gets its own script and its own folder. Anything printed to a screen and copied by hand is not an output; it is a rumour.
Record the parameters that shaped it next to the file: date range, geographic boundary, thresholds, filters, rounding rules, suppression of small cells, and the geographic boundary vintage if boundary shapes changed mid-series. A short block in the project notes covers it:
RUN 2026-10-03 — share of listings up more than 15% YoY
Source: quarterly register extract, snapshot 2026-10-02
Period: 2022Q1 to 2025Q4
Geography: 2024 boundaries, ONS BUA names
Filters: residential only; >=200 listings per area-quarter
Suppression: areas with fewer than 30 qualifying listings hidden
Rounding: one decimal place
Code: commit 4f1c9ab
For published work, record the commit hash alongside the figure. When a correction is needed three weeks later, that hash tells you which version of the code produced the number on screen.
How do you know it worked? The final file carries its own parameters, so nobody has to guess how it was made.
7. Test the Workflow From a Clean State
This is the part people skip, and it is the part that proves everything else. Clone the project into a clean directory on a machine that has never run it, install from the dependency file, run the single entry point, and compare the outputs with what you published.
A Makefile is the cheapest way to get one entry point:
# Makefile
.PHONY: all clean
all: data/output/share_above_15pc.csv
data/interim/clean_rents.csv: scripts/02_clean.py data/raw/rents.csv
python scripts/02_clean.py
data/output/share_above_15pc.csv: scripts/03_analyse.py data/interim/clean_rents.csv
python scripts/03_analyse.py
clean:
rm -rf data/interim/*
Then run the test:
git clone <repo-url> reproducibility-test
cd reproducibility-test
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
make all
diff data/output/share_above_15pc.csv published/share_above_15pc_published.csv
If the diff is empty, your analysis is reproducible. If it is not, do not quietly fix the difference and move on; write down what differed, why, and which version is right. Teams that run this check report it as the moment they discover they were never reproducible in the first place.
Running the script again in your existing working copy proves nothing, because that copy has state, files and installed packages you forgot about. The cold clone is the only real test.
How do you know it worked? A colleague, or you on a borrowed laptop, produces a byte-identical output file without asking you a single question.
8. Package the Analysis for Another Person
Write the README for a stranger who is busy and slightly annoyed. Six sections, no more: the question, the data and its source, the folder structure, how to set up, the order to run things, and the known limitations.
# Share of private listings up more than 15% year over year
Question: How many rental listings rose more than 15% YoY in 2025?
Data: Quarterly register extract, public release.
Snapshot saved 2026-10-02. Licence allows redistribution.
Setup: python -m venv .venv
pip install -r requirements.txt
Run: make all # rebuilds everything
make clean # wipes intermediate files
Output: data/output/share_above_15pc.csv
data/output/chart_yoy.png
Notes: Areas with fewer than 200 listings per quarter are excluded
because year-over-year change is unstable at that size.
Boundaries changed in 2024; the 2024 vintage is used for the
whole series, so area totals are not directly comparable with
the pre-2024 release.
Be honest about limits. Knowing that an area was dropped for low sample size, or that a scrape stopped working when the site changed its markup, saves the next person an afternoon.
If you hand a project to a colleague or an AI coding assistant, put those conventions in a file at the repository root — an AGENTS.md or CLAUDE.md covering folder layout, how to run the pipeline, and the rules for touching data. Tools in this space increasingly expect it, and it costs ten minutes.
How do you know it worked? Someone who has never seen the project can rebuild it using only the README.
Common Mistakes
These are the failures that quietly destroy reproducibility, in rough order of how much damage they cause.
- Undocumented manual edits. Someone fixes a value in the spreadsheet by hand and never records it. Correction: any manual edit becomes a documented, scripted step, with the reason in a comment.
- Editing the source file in place. The downloaded CSV becomes the working file, and the original is gone. Correction: raw folder is read-only by convention, and the cleaning script is the only writer of derived files.
- Unexplained downloads. The data came from somewhere and nobody knows where. Correction: record source, date and licence at the moment of retrieval, not from memory later.
- Hidden filters and manual exclusions. Rows disappear in Excel with no record. Correction: filtering happens in code, and the script prints before-and-after counts.
- Stale or unpinned dependencies. “Works on my machine” is the leading cause of dead pipelines. Correction: commit a dependency or lock file and test the clean install.
- Notebooks executed out of order. The published number depends on hidden kernel state. Correction: exploration in notebooks, pipeline of record in ordered scripts.
- Trusting memory. The person who built the project is also the only person who can rebuild it. Correction: write the README while the details are fresh, then let someone else run it.
- Treating a polished chart as documentation. A good-looking graphic hides the steps that produced it. Correction: the parameters block and the commit hash travel with the output.
Four habits keep all of this standing up. Script the fetch, not just the maths. Print counts and checksums at every stage so problems surface as numbers. Keep exploratory work on a branch so the main pipeline stays clean. And re-run the cold clone whenever you publish, because a pipeline that was never tested is a pipeline you only believe in.
Frequently Asked Questions
What does it mean for an analysis to be reproducible?
It means someone who was not involved in your work can obtain the project, run one command, and produce the same numbers, tables and figures you reported. Not someone who asks you for clarification, and not you on your own machine with three years of context. The test is deliberately strict: it catches undocumented manual steps, unpinned package versions and hidden notebook state, which are the three things that break published analyses most often.
Do I need a Jupyter notebook or a reproducible manuscript to make my analysis reproducible?
No. Notebooks and document-based reports such as Quarto or RMarkdown are useful because prose and code sit together, but they are not required. What matters is that every output is generated by a committed script in a fixed order, that dependencies are pinned, and that a clean environment can rebuild the result. A folder of ordered Python scripts with a Makefile is fully reproducible on its own.
How do I stay reproducible when the data is embargoed, proprietary or cannot be shared?
Keep the retrieval scripted and document the shape of the data rather than the data itself. Commit the fetch script, a data dictionary describing every column, the retrieval date and the licence terms, plus a checksum of the file you used. Readers can then verify the method without seeing the numbers. Where you can share an aggregated or synthetic version, that is even better, because it lets someone else run the pipeline end to end.
How much documentation is enough for a reproducible analysis?
Enough that a competent stranger can rebuild the project without emailing you. In practice that means a README with the research question, the data source and date, the folder structure, setup commands, execution order, expected outputs and known limitations. Parameter notes attached to each output file cover the rest. Length matters far less than whether the reader ever has to guess.
Must a reproducible analysis produce identical numbers every time?
Numerical identity is not required, and small differences from library updates are normal and fine. What must hold is that the outputs agree within the tolerance you state and that any deviation is explainable. Set random seeds once so sampling steps repeat, record rounding rules, and compare results with a stated tolerance rather than demanding bit-for-bit equality you cannot guarantee.
How do I check whether an analysis is really reproducible before I rely on it?
Clone the repository into a new directory on a machine that has never run it, install the dependencies from the committed file, run the single entry point such as make all, and compare the generated outputs with the published figures. If it fails, read the error rather than patching around it, because each failure points at one missing piece of documentation. Never ask the original author to fix it for you.
Conclusion
Start with the next analysis you run, not the last one. Write down the question and the expected result before you compute anything, keep the source file untouched, and push the scripts behind a single entry point before you move on.
Reproducibility is a habit, not a formatting step you do at the end. Projects that survive are the ones where the folder structure, the pinned versions and the README were built as the work happened, one commit at a time.
So write one README this week and run one cold clone before the next thing you publish. If it fails, the error message tells you exactly what to document, and you have just saved yourself the same debugging session six months from now.


