How to Scrape Responsibly and Legally (October 2026)

Collecting pages that anyone can see without logging in is generally lawful, and knowing how to scrape responsibly and legally comes down to four things: what data you take, how you reach it, where you and the site are based, and what you do with it afterwards. Reading public pages at a polite pace is usually fine. Getting past a login, a paywall, or a technical block is a different matter entirely.

This guide gives journalists, developers and newsroom teams a workflow that holds up when an editor, a compliance officer or a court asks why you did it. It is general information, not legal advice. Rules differ by country and state and they change, so treat this as a checklist to run past your own counsel, not a substitute for them.

What You Need

What You Need

You need eight things before a scraper runs, and four of them are documents rather than software. Teams that skip the paperwork are the ones that get letters.

A documented access basis. Either written permission, a published API agreement with terms you accept, or a recorded lawful basis for collecting publicly available material. A vague intention is not a basis.

The site’s own rules. Read the terms of use, robots.txt and any copyright or licensing notice before you read anything else on the site. Save copies with a date, because they change.

A one-page purpose statement. State the question you are answering, who benefits, and what decision the data supports. This single page resolves most internal arguments later.

A privacy read. If the fields you plan to collect can identify a person, you need a named lawful basis under GDPR or an equivalent under CCPA and CPRA, plus a retention date.

An HTTP client with rate limiting built in. Any library that can set a per-host delay, honour a retry budget and identify itself honestly will do. Sophistication is not the point.

Storage with access controls and a deletion date. A scraped CSV sitting in a shared drive forever is a liability, not an archive.

A provenance log. One row per field group recording where it came from, when it was collected, under which permission, and what it was used for. Provenance is what lets you answer a correction request in minutes.

A named owner. One person accountable for the scraper when it breaks, when a publisher objects, and when a deletion request arrives. Unattended scripts need a human on the other end.

How to Scrape Responsibly and Legally: Step-by-Step

How to Scrape Responsibly and Legally: Step-by-Step

How to scrape responsibly and legally before collecting data

Start with the question, not the website. Write down the newsroom or research question in one sentence, then list the fields that answer it and nothing more.

Distinguish three categories before you write a line of code. Public records, such as company filings, court dockets and council minutes, are the cleanest case. Published creative work is copyrightable, so collecting it for analysis is a different act from copying it into your article. Personal data, anything that identifies a person, carries privacy obligations that public availability does not erase.

Then write the intended use in plain language. Aggregate trend reporting, a named investigative finding, and rebuilding a database to sell are three very different projects with three very different risk profiles, even when the fields collected are identical.

You have done this step properly when someone who has never met you can read the purpose statement and predict, roughly, what you will end up with.

Check whether the site allows the intended collection

Terms of use, robots.txt, copyright notices and technical access controls all tell you something, and none of them is the whole answer. Read them in that order.

robots.txt is a voluntary standard, not a law. It is a text file at the root of a site that tells automated crawlers which paths they may and may not fetch, and sites typically include a crawl-delay directive. Treat a disallow as a clear statement of the owner’s wishes. Ignoring it is not automatically a crime, but it removes any argument that you were unaware, and it is routinely the first fact cited in a complaint.

User-agent: newsroom-monitor
Disallow: /account/
Disallow: /search?
Crawl-delay: 10
Allow: /news/

Terms of service become contract questions once you create an account or click accept. A browse-wrap agreement signed by a company employee can bind that company, which is why logged-out collection and logged-in collection are analysed differently. Meta v. Bright Data in the Northern District of California in 2024 turned on precisely that distinction.

If a site publishes an API, a data request address, or a press contact who answers reasonable requests, use it. If the rules are ambiguous, ask in writing and keep the reply. Written silence is not permission, so do not treat an unanswered email as a green light.

Access levelCFAA exposureContract exposurePractical read
Logged-out public pagesLowLowGenerally defensible with rate limits and attribution
Pages behind your own accountLow to moderateHigh once you accepted termsWork from an account whose terms permit your use
Pages behind fake or purchased accountsModerate to highHighAvoid. This is where hiQ v. LinkedIn turned in 2022
Pages behind a paywall, CAPTCHA or IP blockHighHighDo not attempt to defeat the barrier

The gates-up-or-down test is the shorthand US courts use for the CFAA. Van Buren v. United States in 2021 narrowed the law to improper-purpose violations of a use restriction, and hiQ v. LinkedIn in 2022 held that scraping public pages is not access without authorisation. That hiQ still lost, on contract and conduct grounds rather than on the CFAA, is the detail most summaries leave out.

Inspect the site and choose the least intrusive method

Open the target yourself before automating anything. Check whether the same information arrives as a downloadable file, a public API, a sitemap or an open data portal, because an official route removes most of the legal surface at zero cost.

Then sample by hand. Take twenty representative pages, look at how pagination works, note whether content loads after a delay, and record which fields you actually need. This sample is what tells you whether your scraper is going to make fifty requests or fifty thousand.

Prefer plain HTTP requests over a headless browser. A headless browser executes the site’s JavaScript and its anti-bot code at the same time, which is how a modest crawler turns into a load. If pages render client-side, look for the underlying data endpoint or ask the publisher for access.

Never work around authentication, a CAPTCHA, a disabled account or an IP block. Treat every one of those as a decision the owner has already made. If your collection cannot proceed without defeating a barrier, the answer is a data request, not a better crawler.

Record anything you are unsure about while you inspect. Ambiguity belongs in the file now, not in a courtroom later.

Collect only what is necessary and protect people

Field minimisation is the principle that does the most work. Every column you do not collect is one fewer privacy obligation, one fewer copyright question and one fewer thing to delete later.

Hash or drop identifiers you need only to deduplicate. Keep contact details out of the working set and store them separately, encrypted, with named access. Set a retention date on day one rather than deciding in year three whether to keep an archive of people’s home addresses.

Decide in advance whether you will publish individual records or verified aggregates. Publishing a single person’s record from a public register exposes them to harassment regardless of whether the collection was lawful. Publishing a pattern, such as how many permits were issued in one neighbourhood over six months, informs readers without that cost.

Data categoryLegal riskEthical weightNewsroom example
Public recordsLowLow to moderateCompany filings, licensing registers, planning notices
Published news and reference contentModerate, copyright appliesModerateQuoting headlines in an analysis; bulk copying article text
Personal data on public profilesModerate to high under GDPR and CCPAHighAggregating professional profiles into a database
Sensitive categoriesHighHighestHousing, health, employment and immigration status records

Sensitive categories deserve a second sign-off even when the collection is clearly lawful. Housing and health records turn a dataset into something people can be harmed by, so aggregate early and publish late.

Request data at a considerate rate and handle the results

Set one request per few seconds per host, respect any crawl-delay directive, and back off on HTTP 429 rather than retrying harder. Send a user-agent string that identifies your project and includes a contact address, because honest crawlers get treated better than anonymous ones.

Cache what you fetch. Most repeat requests are avoidable, and a cache is the cheapest courtesy you can offer a small publisher.

Give retries a hard ceiling. A scraper that retries forever against a struggling server is the behaviour that turns a technical block into a complaint, and it is the behaviour most likely to show up in a story about a site going down.

Before any analysis, run the verification pass. Match the dataset against the documented purpose, confirm no unnecessary fields survived, spot-check a sample against the live page, and ask whether the collection systematically missed anything. Store that log with the data.

If you use a proxy network, treat the network’s provenance as part of your compliance record. An opt-in residential network and a hijacked-device network can fetch identical pages, but only one of them is defensible to describe.

Common Mistakes

Scraping before reading the terms. A five-minute read of the terms of use and robots.txt prevents most of the problems on this list.

Treating robots.txt as the whole legal answer. It is a statement of preference, not a permission slip, and following it does not remove contract, copyright or privacy exposure.

Collecting everything because storage is cheap. Wide collection multiplies every downstream obligation. Collect the minimum and widen only if the question actually changes.

Ignoring technical restrictions. A CAPTCHA, a blocklist and a disabled account are refusals. The fix is a data request or an API, not a headless browser with more headers.

Hammering a small site. If a local news site goes down during your run, you damaged a real business to answer a marginal question. Rate limits and caching prevent it.

Assuming a public page is fine for any use. Public availability is a fact about access, not a licence about purpose. Republishing scraped text is where copyright exposure really lands.

Publishing personal details without weighing the harm. Verify the pattern, redact identifiers, publish aggregates. The legal permission to collect is not an obligation to publish.

Failing to document sources. Without a provenance log you cannot correct an error, answer a deletion request or defend the story. Build the log with the scraper, not after publication.

When something goes wrong, the sequence is predictable. First an IP block or a 429, then an account termination if you had one, then a cease-and-desist letter, and only then a lawsuit. Practitioners report the first two constantly and the last two rarely. Read the letter, stop collection the same day, work out what you hold, and loop in counsel before replying. Do not argue the legal theory yourself in writing.

If a publisher asks you to delete data you already hold, treat it as a live request rather than a nuisance. Confirm receipt, identify every copy including backups and derived datasets, purge what the request covers, and write down what you deleted and when. Keeping a record of compliance is as valuable as the deletion itself.

Frequently Asked Questions

Usually, yes, when you collect publicly available pages without logging in, without defeating a technical barrier, and without republishing copyrighted content. US cases since Van Buren (2021) and hiQ v. LinkedIn (2022) treat reading public pages as lawful access. Contract terms, copyright, database rights and privacy law still apply, and the answer changes completely once you cross a login, a paywall or a CAPTCHA.

What does robots.txt mean for web scraping?

robots.txt is a voluntary file that tells crawlers which paths to avoid and how long to wait between requests. It is a standard, not a statute, so ignoring it is not automatically illegal. Treat it as a clear signal of the owner’s intent. Following it does not by itself make collection lawful, and courts and complaint letters still cite it as evidence of what you knew.

Should I request an official API instead?

Yes, whenever one exists. An API comes with published terms, agreed rate limits and a named contact, which removes most contract and etiquette risk from a collection. It also gives you structured data instead of parsing work. Ask early, because publishers are more willing to say yes before you build something than after you point a crawler at their site.

Can I scrape personal information for a news investigation?

It can be lawful, but you need a lawful basis under GDPR or an equivalent under CCPA and CPRA, a documented purpose, minimised fields and a retention limit. The stronger constraint is usually editorial. Verify the pattern before you name anyone, aggregate wherever the story allows, and be ready to explain to a person why their record is in your database.

How often should I request pages when scraping?

Slowly enough that the target site feels no load, which in practice means seconds between requests to a single host, slower for small publishers. Respect any crawl-delay directive, cache responses so you never refetch a page you already hold, and back off immediately on HTTP 429 instead of retrying. Nobody has ever been blocked for being too slow.

No, but copying content can be. Collecting a page is generally not a reproduction problem for your own analysis, while republishing scraped text or images as your own typically is. The EU Database Directive adds a separate sui generis right on substantial database investment, which is why Craiglist v. 3Taps and Ryanair v. PR Aviation remain useful references. Link and attribute rather than republish.

Start with the smallest useful version. Write the one-sentence question, pull only the fields that answer it, read the target’s terms and robots.txt before the first request, and log where every value came from.

Knowing how to scrape responsibly and legally is less about clever technique than about restraint: collect less, ask when the rules are unclear, and treat every gate on the site as a decision someone already made. The rest of the machinery, rate limits, user agents, caches and provenance logs, is what proves you did it that way.

Leave a Comment