How to Credit Data Sources Properly: A Newsroom Guide (2026)

Crediting a data source means naming the person or organization that created the data, linking to the original dataset or page, and noting the version or date so a reader can find and check the same numbers you used. In practice that means one source record, built once and reused in every chart, social post and methods note you publish.

Most credit mistakes are not deliberate. A figure travels from an agency to a research paper to a blog to a social thread, and by the time it lands on your desk the closest thing to a source is a screenshot. Learning how to credit data sources properly is mostly about stopping that chain before you quote from it.

The good news is that the work is small. Five minutes per source record, ten if the data is transformed. Once the record exists, the source line writes itself and anyone on the desk can reproduce it.

What You Need

What You Need

Gather these before you draft a single word of attribution. Most broken credits fail because the publisher name or the release date was never written down while the researcher was still on the phone.

  • Source name — the title of the dataset, table or series exactly as the publisher writes it.
  • Publisher — the organization that released it, plus the individual author or team when there is one.
  • Release or update date — and the version number if the publisher keeps one.
  • Access date — the day you pulled the file, which matters more for live databases than for PDFs.
  • Direct URL — the canonical page or file, not a search result or a shortened link.
  • Table or record reference — the specific table, sheet, endpoint or query that holds the number.
  • Units, definitions and coverage — currency, population base, geographic scope, and any exclusions the publisher documents.
  • Licence or terms of use — the reuse conditions attached to the data, copied from the publisher’s own page.
  • Methodology note — what you filtered, joined, weighted or excluded yourself.

A spreadsheet row per source is enough. So is a note in a research notebook, as long as the same fields appear every time. Tools like Zotero or a shared Airtable base help, but the fields matter more than the tool.

Step-by-Step: How to Credit Data Sources Properly

Six steps, in order, take a raw number from an unnamed chart to a verifiable credit. The test of each step is simple: could a colleague open what you saved and land on the same figure you published?

1. Identify the original source

The single most useful habit in data attribution is chasing a number back until it stops moving. Publishers get layered: an agency releases a table, a journal summarizes it, a blog quotes the journal, and a social post quotes the blog with the number rounded.

Work backwards from the quote. If you only have a chart image, look for the underlying table or dataset. If the text names an agency, find that agency’s release rather than the article that mentioned it. If the number came from an API response, note the endpoint and parameters rather than the platform’s marketing name.

It worked when you can name the body that produced the observation itself, not the body that described it. Keep the chain anyway: “primary” tells readers who owns the evidence, and naming the intermediary separately keeps the trail honest.

2. Verify the dataset and publication details

Open the original page and record what it actually says. Dataset title, publisher, release date, last-updated date, version, geographic coverage, units and definitions all belong in the record. Government statistics pages in particular often carry a revision note; the figure you ran may not be the figure now published.

Check the definitions before the numbers. A rise in “cases” that quietly changed definition in the middle of the series is a different story, and the credit should not hide that. Where a data dictionary exists, save a copy alongside the file.

3. Write a complete attribution

A usable newsroom credit follows a fixed formula: publisher, dataset, version or year, table reference, access date, link. Keep it under about fifteen words so it fits under a graphic. Knowing how to credit data sources properly comes down to using the same fields in the same order every single time, so a reader learns what to expect after the first credit they read.

Templates you can copy:

  • Source: PUBLISHER (YEAR), table TREF
  • Data from: PORTAL, DATASET NAME, vVERSION, accessed DATE
  • Data from API NAME (ENDPOINT), retrieved DATE; analysis by OUTLET
  • Source: SURVEY NAME by RESEARCH ORG, fieldwork MONTH-YEAR, n=FIGURE
  • Data from X, analysis by Y — the split that credits the data maker and the analyst separately.

Worked examples: a national statistics office table becomes Source: National Statistics Office, Household Expenditure Survey, table 4.2. An open data portal download becomes Data from Open Data Portal, Business Demographics, v2.1, accessed 14 March 2026. A public API becomes Data from Census Data API (endpoint: population by age band), retrieved 2 June 2026. A commissioned survey becomes Source: Regional Health Survey by Civic Research Group, fieldwork May 2026, n=4,200.

Creative Commons licences add one clause: name the licence and link it. Under CC BY you must credit the creator, indicate the licence, link to the licence, and link to the original material. CC BY-SA adds a share-alike condition on adaptations. CC BY-NC and CC BY-ND restrict commercial use or derivatives, which newsrooms and news apps hit immediately. CC0 waives rights but asks for attribution anyway.

Put the link where the reader can reach it in one click: in the source line beneath the graphic, in a footnote on mobile, and in full in the story’s methods note. On social, the link goes in the post itself, not in a reply.

Deep-link where you can. A table anchor beats a landing page, and a specific record beats a table of fifty pages. When no stable link exists, say so and name the access date instead, so a reader knows the state of the data you saw.

Then disclose what you did to it. Filters, exclusions, joins, reweighting and unit conversions all change the meaning of a number. Say it in plain words: “Records without a postcode were excluded, about 4% of the file” is worth more to a reader than a methods appendix they will never open.

5. Add access and reuse information when relevant

Add an access date whenever the data changes under you: live dashboards, API endpoints, crowdsourced counts and anything that gets revised. PDFs with an edition and a fixed release date usually do not need one, though nobody is harmed by including it.

Record the file format and version too. “v2.1” or “the 2026 edition” tells a reader whether they are looking at the same artifact you used. Where a dataset has a DOI or another persistent identifier from DataCite, use it in preference to a URL that may move; the identifier keeps resolving even when the hosting page is reorganized.

Be clear with yourself about what a credit is and is not. Attribution is not permission, and a licence that permits reuse may still require you to seek permission for something specific, such as republishing another outlet’s chart or a database dump. When the terms are unclear or the use is unusual, that is a conversation for an editor or a lawyer, not a judgement call at the keyboard.

6. Store the credit with the story

Write the record somewhere durable, in the same place as the analysis: the CMS entry for the story, the project README, a notebook cell, or a shared source log. Give each record an identifier so the graphic, the story text and the newsletter all pull from one entry.

A workable record looks like this:

  • id, publisher, dataset_title, creator, version, release_date
  • accessed, canonical_url, table_ref, units, coverage
  • transformations, limitations, licence, attribution_text

You have done the job when a colleague can regenerate the credit, an editor can check the link six months later, and you can update one field when the publisher revises the data. Machine-readable formats help: a comment header at the top of a script, a data block in a README, or metadata in a notebook all cost a minute and save an argument later.

Crediting the people who made the data

Once the record exists, sending it back to the source costs a short email and pays off more often than not. Tell the researcher what you published, link to it, and ask whether the credit reads correctly to them. That single message is sometimes the difference between a source who answers your next question and one who does not.

Ask for co-authorship when the analysis genuinely depends on someone’s expertise, and offer it before you ask for a quote. Practitioners describe the relationship in plain terms: using a colleague’s months of fieldwork without a mention is closer to borrowing than to citing.

Common Mistakes

These are the errors that show up repeatedly in published work, each with the fix I would apply.

  1. Crediting the aggregator instead of the originator. Fix: one more backwards search. Find the agency, repository or research group that produced the numbers and name them first; list intermediaries separately.
  2. Linking to a home page or a search result. Fix: link to the specific table, file or record, and re-check the link before you publish.
  3. Omitting the date, so the credit cannot be checked. Fix: record the release date and, for live data, the access date.
  4. Ignoring the version. Fix: store the version or edition string exactly as the publisher writes it.
  5. Failing to describe transformations. Fix: add one line covering filters, joins, exclusions and unit conversions, with rough proportions where they matter.
  6. Treating a database download as a stable source. Fix: note that it is a snapshot, with the access date, and link to the live dataset too.
  7. Using a broken or shortened URL. Fix: canonical links only, checked after publication.
  8. Confusing credit with permission. Fix: read the licence terms, and ask an editor when the intended use is unusual.
  9. Losing the source inside AI-generated summaries. Fix: keep the machine-readable record next to the analysis, and never let a model be the only place the citation exists.
  10. Crediting nobody for first-party data. Fix: name the team that collected the data, including survey respondents in aggregate terms.

A final habit catches most of these: read the source line out loud before the piece goes live. If it sounds vague, it is vague. If a reader could not tell what you did to the data, tell them.

Frequently Asked Questions

How do I cite data sources?

Name the publisher, the dataset title, the year or version, the specific table or record, and a direct link, plus the access date for anything that changes. Put that credit in a source line under the graphic or in a methods note, and add one sentence describing anything you changed. For a database with a DOI, use the DOI so the link keeps resolving after a redesign.

What is the difference between citation, attribution, acknowledgement and licensing?

Citation points at the source in a reference list so a reader can find it. Attribution credits the creator in published text or a source line. Acknowledgement thanks someone who contributed but is not a source of evidence. Licensing sets the legal conditions for reuse, such as attribution and share-alike under Creative Commons. One source record can produce all four, but only the licence grants rights.

Do I have to credit data that is in the public domain?

Legally, most public-domain material requires no credit. Practically, you still should, because readers use the credit to judge the data and because the creator may not have released everything. Treat it as professional courtesy rather than an obligation, keep the credit short, and never imply the data is endorsed by whoever published it.

What counts as the primary source for a statistic?

The primary source is the body that produced the underlying observation, such as the statistics agency, research group or organisation that collected the records. Articles, posts and summaries are secondary sources, however widely they are shared. Credit the primary source first, and name the secondary source separately when it added context or a different reading of the figures.

How do I credit a dataset used in a chart, map or interactive graphic?

Place a source line directly beneath the graphic with the publisher, dataset, year or version and a direct link. Repeat the credit in the surrounding story text so it survives screenshots and embeds. In an interactive, keep the credit visible rather than hidden behind an info icon, and put the full methods note in a linked panel that lists your filters and calculations.

How should I credit data used in AI-generated analysis?

Credit the underlying dataset, not the model. Record the source, version and access date in the project file before you run any analysis, then keep that record out of the chat window so it cannot be lost. Say plainly when a language model generated text or code from those inputs, and never let a model’s summary be the only surviving trace of where the numbers came from.

Conclusion

Start with one source record. Put the original publisher, the dataset or page, the version or release date, the access date, the direct URL, the transformations you made and the attribution text in one entry, then reuse that same entry across the story, the chart, the social post and the project documentation.

Everything else in this guide is a variation on that habit. Once the record exists, crediting your sources stops being a favour you do at the end and becomes a by-product of the work itself.

Leave a Comment