Using API data for a story means pulling records straight from a machine-readable endpoint and turning them into a finding somebody else can check. The workflow barely changes from one project to the next: define the reporting question, find the source, request the data, document the pull, clean and verify it, then publish it with a methods note.
Last updated: October 2026.
Table of Contents
- What You Need
- Step-by-Step
- Common Mistakes
- Frequently Asked Questions
- Do I need coding skills to use API data for a story?
- How can I download and analyze a large API dataset?
- How do I verify information that comes from an API?
- What should I do when an API returns missing values?
- How should I cite an API and the data retrieved from it?
- Can API data prove that one thing caused another?
What You Need

Most of this is preparation, not software. Before you touch an endpoint, sort out these seven things.
- A reporting question specific enough that data could actually answer it.
- A source that publishes an API or a bulk download, plus its documentation.
- Credentials if the source requires them: an account, an API key, a working email address.
- A place to keep the records: a spreadsheet for small pulls, OpenRefine, a SQL database or a notebook for anything larger.
- A text editor and a saved copy of every raw response you receive.
- A data dictionary: one row per field, with its meaning, type, units and source.
- Basic fact-checking habits and a second person who will rerun your pull before publication.
Coding experience helps, but it is not the gate. A spreadsheet formula, a browser console or a no-code scraper will get you the same records for a first look, and the scripts come later once the manual work becomes the bottleneck.
Step-by-Step
Eight steps, in the order that avoids wasted evenings. The two that catch newcomers are step 3 and step 6.
1. Start with a reporting question
Turn a vague interest into a question the dataset can settle. Write down the population you care about, the time period, the geography, the comparison you want to make and the decision the reporting could inform. “Air quality” is not a question; “which neighborhoods saw the largest increase in fine-particle readings between two years” is.
The most common mistake here is letting the fields decide the story. If the endpoint exposes a county_fips column, that is not a reason to write a county story. Start from what you suspect, then check the data can support it.
2. Find and assess the right API
Search government portals, statistical agencies, regulators and institutional publishers before you search anything else. Then read the documentation properly, not just the quickstart: what the source covers, how often it updates, which variables exist, the licence, the access limits and the methodology notes.
Distinguish an official endpoint from something you discovered by watching network traffic on a website. An undocumented call behind a public page is not a supported API, and using it programmatically carries the same risk as scraping. Record the API name, version, base URL and the date you retrieved the documentation.
3. Make a small test request
Make one small request before attempting a large download, and save the response as a reference file. The same call can be made three ways, and all three return identical records.
| Path | Best for | Watch out for |
|---|---|---|
| Browser developer tools console or a hosted request builder | One-off exploration, learning the shape of the data | Manual, and easy to lose track of what you requested |
| Spreadsheet function such as IMPORTDATA, pointed at a URL that returns a flat file | Flat CSV or JSON endpoints, quick sorting and filtering | No authentication, no pagination, large pulls time out |
| A short Python or JavaScript script | Keys, pagination, thousands of records, repeatable refreshes | Slower to start; needs a text editor and a terminal |
While testing, learn what the provider does with authentication, query parameters, pagination, rate limits and error responses. A 200 response with zero results is not a broken key, and a 429 means you asked too fast, not that you did something wrong.
4. Inspect and document the data
Open the response and identify the format, then find one complete record and walk through it field by field. Note which fields are required, which are optional, which values are nested objects or lists, which observations are missing, and what the units and codes mean.
Build the data dictionary here, before cleaning. Watch for fields that look numeric but are not measurements: identifiers, coded categories, percentages stored as whole numbers, or a status flag where 1 and 0 mean something other than true and false. Keep the original JSON or CSV untouched and work on a copy.
5. Clean and validate the records
Use a fixed sequence, and write down what you changed. Remove duplicate identifiers. Standardise dates into one format and place names into one spelling. Convert units consistently. Preserve missing values as missing rather than filling them with zero. Check for impossible values such as negative counts or percentages above 100. Reconcile your row count and totals against the totals the provider publishes.
A spreadsheet is fine up to a few thousand tidy rows. Past that, OpenRefine handles messy records well, SQL handles joins and group-by questions, and Python or R pay off once you want the same checks rerun every week. Whichever you use, the changes must be reproducible by someone who was not in the room.
6. Find a defensible pattern
This is the step the tutorials usually skip, and it is the one that decides whether you have a story. Sort and group the data. Compare groups. Look for change over time and for differences between places. Then convert counts to rates whenever the populations behind them differ in size, and say which denominator you used.
As an illustration, suppose a city reports 400 inspections in one district and 120 in another. On its own that says nothing, because the districts differ in size. Divide by residents or by number of licensed properties and the ordering can change entirely. That reordering is the story.
Then test it. Does the pattern survive dropping the largest single record? Does it hold under a slightly different definition of the same measure? A finding that depends on one exclusion or one vague category is not yet a finding.
7. Verify, contextualize and disclose
Compare your result with the provider’s own published figures or charts for the same period. Read the methodology notes for revised definitions, restatements and known gaps. Check whether the series has been revised since you pulled it. Seek human context from the agency, from experts or from the people named in the records, and where the number is load-bearing, confirm it against an independent source.
API results are evidence, not automatic truth. Write down the retrieval dates, the filters you applied, the records you excluded and the limitations you know about. That note becomes your methods box.
8. Present the evidence clearly
Pick the form that matches the question: a chart for change over time or comparison between groups, a map for geography, a table when the reader needs to look up exact values, and plain writing when one number carries the point. Label units, dates, geography and source on the graphic itself, explain the denominator in a line of text, and keep the axis starting where it should so a small difference does not look enormous.
Finish with a short methods note: the source, the endpoint, the filters, the retrieval date and what the data cannot tell you. Link to the public dataset if there is one. Readers who disagree with your conclusion should be able to reach the same records you did.
Common Mistakes
- Treating a sample as the whole population. Check your row count against the total the provider publishes before you analyse anything.
- Ignoring pagination. You pulled page one of forty and every number downstream is wrong. Most responses tell you how many pages exist.
- Reading missing values as zeros. An absent field often means not collected or suppressed, and those cases distort averages badly.
- Comparing raw counts across places of different sizes. Convert to rates and state the denominator.
- Mixing incompatible definitions. Two agencies can measure the same thing differently; check the methodology notes before putting both on one chart.
- Failing to record your filters. An unfiltered endpoint is not the same dataset on two different days.
- Describing correlation as cause. Say what changed alongside what, and reserve cause for evidence the dataset does not contain.
- Publishing a graphic with no source note. Source, retrieval date, filters and limitations belong under the chart, not in your notebook.
Three habits prevent nearly all of it. Work from a saved raw file rather than a live endpoint, script the checks that would catch a silent truncation, and have one colleague rerun your request before the piece is edited. If they cannot reproduce it, neither can your editor.
Frequently Asked Questions
Do I need coding skills to use API data for a story?
Not to start. Most public-reporting APIs return data you can pull through a spreadsheet function, a browser console or a no-code scraper, then sort and filter by hand. Coding earns its keep when the dataset is large, arrives nested, needs thousands of pages, or has to refresh on a schedule. Learn the request step with a tool first, then pick up one script language when the manual work becomes the bottleneck.
How can I download and analyze a large API dataset?
Use the provider’s bulk download or full-dump file when one exists, since it beats thousands of separate requests. Otherwise loop through the pagination, write each page to disk as you go, and respect the rate limit by pausing between calls. Then load everything into a tool built for the size: OpenRefine for cleaning, SQL for querying, a notebook for reproducible analysis. Never try to hold a large pull in a browser tab.
How do I verify information that comes from an API?
Compare your numbers with the provider’s own published totals or charts, then read the methodology notes for definitions, revisions and known gaps. Record the endpoint, parameters and retrieval date so a colleague can rerun the request and get the same result. Where a figure is load-bearing, confirm it against an independent source or from the people named in the records before it goes to print.
What should I do when an API returns missing values?
Do not convert blanks to zero. A missing field usually means not collected, not applicable, or suppressed, and each means something different for your analysis. Keep missing values visible in the working file, count how many records are affected, and check whether the gap clusters in one place, period or group. If a suppression rule hides small counts, say so in the methods note rather than dropping the rows quietly.
How should I cite an API and the data retrieved from it?
Name the organisation publishing the data, the dataset or endpoint, the date you retrieved it and the version where one applies. APIs change without notice, so a number pulled in March may not match today’s response. A workable methods note lists the endpoint, the filters applied, the retrieval date and any exclusions, which lets a colleague reproduce the pull and a reader see its limits.
Can API data prove that one thing caused another?
No. An API records what was measured, not why. A pattern across thousands of records is strong evidence that two things move together, but the same numbers appear for reasons the dataset never captures. Arguing cause needs a mechanism, a comparison or evidence from outside the data, such as interviews, internal documents or prior research. Until you have that, write the finding as an association.
If you take one action from this, write the reporting question before you open a single endpoint, then check whether the source you have in mind actually publishes it. A well-posed question on a well-documented source takes an afternoon; a vague one on an undocumented endpoint takes a fortnight and usually ends in a correction.


