To download data from a web portal in bulk, you pull many records or files in one automated pass instead of saving them by hand. There are three real routes, and you should try them in order: the portal’s own export button, its documented API, then a scripted client such as wget, curl or Python. Check the first two before automating the third, and confirm bulk access is permitted before you start either.
Most newsroom and research teams outgrow the export button quickly. A manual export usually caps at a few hundred rows, truncates without warning, or times out on a large filter, which is why a silent partial download is the most common failure in this workflow.
By bulk, most people mean more than about 100 items or more than 50 MB in a single request. Below that, a manual download is fine and faster. Above it, you need a client that can resume, retry and prove what arrived, and a written record of where the data came from so the next person can repeat the pull.
Whatever route you pick, the work splits into the same four jobs: work out what kind of portal you are dealing with, get permission, move the data with a client matched to the portal type, then validate and document the result. The rest of this guide walks through those jobs in order.
Table of Contents
- What You Need
- Step-by-Step: How to Download Data from a Web Portal in Bulk
- 1. Check the portal’s export, API, and bulk-access options
- 2. Confirm that bulk downloading is permitted
- 3. Choose CSV, JSON, database, or archive format
- 4. Use the portal’s native bulk export when available
- 5. Automate an API download from the web portal in bulk with pagination and rate limits
- 6. Split large files and preserve the original source data
- 7. Validate row counts, dates, columns, and encoding
- 8. Store, document, and refresh the downloaded dataset
- Common Mistakes
- Frequently Asked Questions
- Conclusion
What You Need

You need four things before the first request goes out: the portal’s own documentation, access you are actually permitted to use, a client that matches the portal type, and somewhere to record what you pulled.
- Portal documentation. The help page, developer docs or data catalogue that says which formats and query parameters exist. Most portals document a bulk path that search results never surface.
- Permitted access. An API key, account, .netrc credential, cookie session or single sign-on login, depending on the portal. Public pages and licensed pages are not the same thing.
- A client. The export screen, an API call, wget, curl, a Python script with requests, or a headless browser for portals that only render through JavaScript.
- Storage with headroom. Estimate the total size first, and keep at least double it free so a half-finished archive cannot fill the disk. NASA Earthdata workflows, for example, start by counting granules and total bytes before any transfer.
- A record. A text file or spreadsheet for the source URL, the exact query parameters, the retrieval date, the licence terms and the portal version. Five fields now saves an argument later.
Step-by-Step: How to Download Data from a Web Portal in Bulk
Portals differ more in interface than in substance. Match the approach to the type you have, and most of the guesswork disappears.
| What the portal looks like | Best approach | Watch for |
|---|---|---|
| Open directory listing of files with a visible index page | wget recursive mirror, or a link list piped to curl | Depth limits so the crawl stays inside the dataset |
| Documented REST or GraphQL API with keys | Scripted API client with pagination and throttling | Page size caps, rate limits, pagination cursors |
| Search form with dropdowns, date ranges and a Search button | Scripted POST that replays the form, or a headless browser | CSRF tokens, session expiry, POST-only result pages |
| Queued job: “your export will be ready in two hours” | Submit the request, then poll for the download link | Expiring links, single-use links, notification email only |
| Published cloud bucket or public dataset | Cloud CLI against the bucket, in-region | Credentials, region and egress cost |
1. Check the portal’s export, API, and bulk-access options
Look for the export button, the API documentation and the data catalogue before you look at any script. If a portal offers a native bulk export, an item basket, or a documented endpoint, that is the intended path and it will be faster and more stable than anything you build yourself.
Check four specific things: whether exports can be filtered by date range or region before they are generated, whether the API documents pagination and a maximum page size, whether rate limits are stated in requests per second or per day, and whether the licence restricts redistribution or automated access.
2. Confirm that bulk downloading is permitted
Automated access is not automatically allowed, even when the data is publicly visible. Read the terms of use, the robots.txt file at the site root, and any API documentation that states access policy, and note the rules that apply to your pull.
Three questions settle most cases: does the portal publish an API or bulk channel for the data you want, does it cap how much one account may pull, and does it ask you to identify yourself with a user-agent string or contact address. Where a data transfer agreement exists, that beats any scripted workaround. When the rules are genuinely unclear and the pull is large, email the portal owner and ask for a data transfer instead of guessing.
3. Choose CSV, JSON, database, or archive format
Pick the format that matches how you will use the data, not the format that downloads first. CSV is the most compatible for analysis but flattens nested records, JSON preserves structure at the cost of parsing work, and a database dump or archive is the right choice for a large relational extract.
| Format | Strength | Weakness | Reach for it when |
|---|---|---|---|
| CSV | Opens everywhere, small, easy to diff between runs | No nesting, delimiter and encoding traps, no types | Flat record sets, spreadsheets, quick analysis |
| JSON or JSON Lines | Keeps nested objects and arrays intact | Verbose, slower to parse, harder to eyeball | Records with nested attributes or lists |
| SQL dump or database file | Preserves keys and relations exactly | Portal-specific, needs a matching engine | Relational extracts that must stay joinable |
| ZIP, TAR or gzip archive | Many files in one transfer, lower overhead | Truncates silently if a transfer breaks midway | Multi-file collections and long-running pulls |
| Parquet or columnar format | Small, fast to query across systems | Not readable without a tool | Large analytical loads feeding a pipeline |
Encoding deserves its own decision. Most portals export UTF-8, and a surprising number still send Latin-1 or add a byte order mark that breaks the first column name. Check the first few rows before you build anything on top of the file.
4. Use the portal’s native bulk export when available
A native export is usually the fastest route, so use it whenever the portal offers one. Apply your filters in the interface first: a date range, a region, a record type, a bounding box. Filtering at the source means you download far less, and it keeps the request small enough to succeed on the first try.
Start the export, note the job reference or confirmation, and then check how delivery works. Some portals email a link when the file is ready, which can take hours, and some hand you a link that expires or works only once. Download the archive as soon as the link appears, test that it opens without errors, and compare the record count inside against the count the portal reported for your filter. That single comparison catches most silent partial exports.
5. Automate an API download from the web portal in bulk with pagination and rate limits
A documented API is the most reliable bulk route because it returns structured records and honours your filters. The workflow is the same whatever language you use: authenticate, request one page, read the pagination token, wait, repeat, and write each page to disk as it arrives.
import time, requests
BASE = "https://portal.example.org/api/v2/records"
session = requests.Session()
session.headers.update({"Authorization": "Bearer YOUR_TOKEN",
"User-Agent": "newsroom-research/1.0 ([email protected])"})
params = {"from": "2026-01-01", "to": "2026-12-31", "limit": 1000, "cursor": None}
page, saved = 0, 0
while True:
r = session.get(BASE, params={k: v for k, v in params.items() if v}, timeout=60)
r.raise_for_status()
body = r.json()
rows = body.get("results", [])
if not rows:
break
page += 1
with open(f"page_{page:04d}.json", "w", encoding="utf-8") as fh:
json.dump(rows, fh, ensure_ascii=False)
saved += len(rows)
params["cursor"] = body.get("next_cursor") # cursor or offset page
time.sleep(1) # throttle; read the docs for the real cap
if not body.get("next_cursor"):
break
print("pages:", page, "records:", saved)
Four details separate a script that works from one that gets blocked. Honour the stated rate limit rather than testing it, use a session so the connection is reused, retry on 429 and 5xx responses with a growing delay instead of failing the run, and name files deterministically so a re-run skips what already exists.
On “load more” buttons and offset pagination, the same principle applies in reverse: find the parameter that the button sets, such as a page number, offset or cursor, and drive it yourself. Never assume page two exists at a guessable URL on a portal that hides it behind a form. If results only appear after a POST, you must replay the POST, which means carrying session cookies and, on most portals, a CSRF token read from the form page before the first submission.
6. Split large files and preserve the original source data
Big pulls fail at predictable moments: a session expires, a token is refreshed, a connection drops, an archive is cut short. Downloading in batches, by date range or by record type, turns one fragile hour-long job into several short ones you can resume.
Keep source files untouched in their own folder, and write anything you change to a separate location. Rename, filter, convert and clean copies, never originals. Where the portal publishes checksums or file sizes, record them at download time so you can prove later that nothing arrived truncated, and keep a manifest listing each file, its size, its retrieval date and the query that produced it.
7. Validate row counts, dates, columns, and encoding
Validation is the step people skip, and it is the step that tells you whether the last two hours were useful. Count rows and compare against the number the portal reported. Check the earliest and latest dates in the extract against the range you requested, since a truncated pull often keeps the most recent records and quietly drops the oldest.
Compare the column list against the previous run, because fields get added, renamed or emptied without notice. Look for duplicate identifiers, and open a sample of non-Latin text to confirm the characters are intact. Finally, confirm the file is complete by testing the archive rather than trusting the file size shown in the browser.
8. Store, document, and refresh the downloaded dataset
Treat the download as the start of a dataset, not a finished file. Store it with a name that carries the portal, the filter and the date, then write a short provenance note: source URL, query parameters, retrieval date, licence, portal version, and any transformation you applied.
Plan the refresh before you need it. A recurring job, an incremental sync, or a monthly manual re-run with the same parameters means you can diff each new file against the previous one and spot a change in the underlying data rather than in your own code. For newsroom work, that diff is often the story: a series that stopped updating, a column that emptied out, a record type that vanished.
Common Mistakes
Almost every failed bulk download is one of the problems below, and each has a straightforward fix.
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 Forbidden on most requests | No valid session, missing user-agent, or terms forbid automation | Authenticate properly, identify your client, re-read the access policy |
| 429 Too Many Requests partway through | Request rate above the published limit | Slow down, add backoff on retries, split the pull into smaller jobs |
| CAPTCHA or a human check appears | Traffic pattern looks automated | Stop and request access, or request a data transfer from the portal owner |
| Run dies partway through | Session or token expired during a long job | Download in batches, refresh authentication, resume from a manifest |
| Archive will not open | Transfer truncated, often with no error message | Test the archive after download, retry with resume enabled, verify the checksum |
| Fewer rows than expected | Pagination stopped early, or the portal capped the export | Log the row count per page, follow the documented cursor, confirm the portal’s row cap |
| Garbled characters in names or addresses | Encoding mismatch, or a byte order mark | Request UTF-8 explicitly, read the file with UTF-8 and strip the mark if present |
| Last month’s file looks like this month’s | Raw files overwritten in place | Keep source files read-only, name by date, diff each refresh against the previous run |
| Published figure is wrong | No validation before analysis, licence ignored | Validate row counts and dates first, and check the licence before publishing |
Two habits prevent most of the above. Slow down before you speed up, since a rate-limited run that retries harder gets blocked harder. And never treat a download as data until you have counted it.
Frequently Asked Questions
How do I download data from a web portal in bulk?
There are three routes, used in this order. First, the portal’s own export or item basket, with filters applied before you generate the file. Second, a documented API, called from a script with pagination, throttling and retries. Third, a scripted client such as wget, curl or Python, used when no export or API exists. Check the portal’s documentation for all three before writing any code.
Is bulk downloading from a website legal?
It depends on the portal, not on the data. Public visibility is not permission. Read the terms of use, robots.txt and any API access policy, and check whether the data is licensed for redistribution. Where the rules are unclear and the pull is large, contact the portal owner or submit a formal information request instead of automating anyway. Record what you agreed to, and keep the record with the data.
How do I download multiple files at once?
Collect the file URLs into a list first, then feed that list to a downloader. wget handles a recursive mirror of a public directory with depth and file-type filters. curl handles a prepared list, one transfer per line, and xargs can run several at once with a concurrency cap you set. Keep the cap low and add a delay between files, because parallel requests are the fastest way to hit a rate limit.
Is there a free bulk file downloader?
Yes, and the free options are usually enough. wget, curl and aria2 ship with most Linux and macOS systems. Python with the requests library handles anything that needs authentication, pagination or logging. Graphical tools such as JDownloader and DownThemAll! cover link-list downloads for people who would rather not open a terminal. Paid tools mostly add scheduling, queues and shared cloud storage.
How do I extract data from a website using Python?
Use requests to fetch the page or endpoint, then parse the response. JSON endpoints need only the json method. HTML pages need a parser such as BeautifulSoup, and pages that render through JavaScript need a headless browser such as Playwright. Build the file list in one pass, then download from that list so parsing and transferring stay separate, each with its own retry logic and log.
How do I download a large dataset without getting blocked?
Slow down before you speed up. Set a request rate below the published limit, keep one session open so connections are reused, retry 429 and 5xx responses with a growing delay, and identify your client with a descriptive user-agent that includes a contact address. Download in batches so a long job can resume, and stop immediately if a CAPTCHA appears rather than trying to work around it.
Conclusion
To download data from a web portal in bulk, do four things in order. Read the portal’s documentation and find out which of the three routes it supports. Confirm that bulk access is permitted, and request a data transfer if the rules are unclear.
Then pick the least disruptive method, download less by filtering at the source, and validate the first file completely before you automate anything. A single verified export beats an unverified script every time, and this guide is current as of October 2026.


