How to organize documents for an investigation means building one folder system where every file can be found in seconds and traced back to who produced it, when it arrived, and which claim it supports. The practical version is a predictable folder tree, a file-naming convention that encodes date and source, and a register that logs each document’s origin.
That system takes about an hour to build and pays for itself the first time someone asks, “where did that spreadsheet come from?” It matters most for journalists running long or collaborative investigations, in-house and freelance investigators, workplace and HR investigators handling misconduct complaints, and legal teams managing discovery.
People ask for the elaborate version first. They end up with forty folders, three competing copies of the same interview notes, and a naming scheme nobody else can decode. The simpler the structure, the more likely it is still being used six months later.
Table of Contents
- What You Need
- Step-by-Step: How to Organize Documents for an Investigation
- Common Mistakes That Break an Investigation File
- Frequently Asked Questions
- What is the best way to organize documents for an investigation?
- Should I keep original and edited investigation files together?
- How should I name files when working with confidential sources?
- What metadata should an investigation document register contain?
- Is cloud storage safe for sensitive investigation documents?
- How often should I back up an investigation folder?
What You Need
Seven things, and none of them require a paid tool. If you have the first four, you can start working today and add the rest as the file grows.
- A working folder. One parent directory that holds everything, with nothing from this investigation stored outside it.
- An investigation brief. A single page stating the reporting question, the scope, the deadline, and the specific claims you are trying to establish.
- A file-naming convention. Written down as a formula, not held in your head.
- An evidence or source register. A spreadsheet where each row is one document, one interview, or one physical item.
- A tracking sheet. A status view built on the register showing what is outstanding, what is verified, and what is waiting on someone else.
- Backups. At least one copy on separate media, disconnected from your working machine.
- Access controls. A written list of who can see the material, at what level.
None of these need to be bought. Local storage, a shared drive, a database table, or a newsroom collaboration platform can all hold the same structure, as long as the folder tree and the naming convention match the register. Where the material lives is a decision about access and search; how it is named and logged is a decision about discipline, and that part does not change between platforms.
Step-by-Step: How to Organize Documents for an Investigation
Define the investigation structure
Build the top-level folder tree around your reporting question rather than around document types. Workstreams, date ranges, sources, locations, and evidence types all work; a single “Documents” folder never does.
A structure that survives contact with a real investigation looks like this:
INV-2026-014_Harbor_Contracts/
00_Brief/— scope, questions, timeline, assignment letters01_Evidence/— one subfolder per evidence ID (EV-001, EV-002)02_Interviews/— one subfolder per person, dated03_Records_Requests/— request letters, responses, delivery receipts04_Data/— raw downloads, cleaned extracts, analysis notebooks05_Public_Records/— filings, permits, published reports06_Drafts/— story drafts and fact-check worksheets07_Source_Protection/— separated, encrypted, restricted08_Admin/— budgets, legal review, publication correspondence
Give the investigation a short code like INV-2026-014 and put it at the top. It appears in the folder name, in every filename, and in the register, which means a single search term pulls the whole project out of whatever drive it lives on.
How to tell it worked: pick any ten files at random and try to assign each one to exactly one location in the tree. If two files compete for the same place, the tree is doing too much. If a file has no home, your scope is looser than your brief says.
Create a consistent file-naming system
A good filename answers four questions without opening the file: when is it from, which investigation is it for, where did it come from, and what is it. A format like YYYY-MM-DD_INV-CODE_SOURCE_DOCTYPE_vNN_STATUS.ext covers all four.
| Segment | Purpose | Example |
|---|---|---|
| Date | Sorts chronologically everywhere; use YYYY-MM-DD so it is never ambiguous | 2026-03-14 |
| Investigation code | Isolates one project from every other project on the drive | INV-2026-014 |
| Source | Tells you who or what produced the material | PortAuthority |
| Document type | Fixed vocabulary, never invented ad hoc | Invoice, Interview, Email |
| Version | Ends the final_FINAL_v2 spiral before it starts | v03 |
| Status | Tells a reader whether the file is safe to cite | VERIFIED, DRAFT, HOLD |
The before-and-after is usually the fastest way to see the difference:
IMG_4471.pdfbecomes2026-03-14_INV-2026-014_PortAuthority_Invoice_v01_HOLD.pdfnotes from dana.docxbecomes2026-04-02_INV-2026-014_DanaOkafor_Interview_v02_VERIFIED.docxFinal FINAL v2 (really final).xlsxbecomes2026-05-09_INV-2026-014_HarborBoard_ContractRegister_v03_VERIFIED.xlsxRivera interview FINAL.docxbecomes2026-04-18_INV-2026-014_MRivera_Interview_v01_VERIFIED.docx
Two rules keep names searchable without exposing anything. First, never put a source’s real identity in a filename if the folder is shared more widely than the source’s protection level allows; use a source code that maps to the name only in the protected register. Second, write dates as text and in full, and skip special characters like slashes, colons, and quotation marks, which break files on some systems.
How to tell it worked: sort the folder by name and you should get a clean chronological reading of the investigation. Then have someone who did not build the system find a document you mention in a sentence. If they cannot, the convention is too clever.
Separate originals from working copies
Originals are the material as received, unedited, unrenamed in content, and never opened for writing. Working copies are everything you derive from them: transcriptions, translations, annotations, redacted versions, spreadsheet extracts.
Keep them in separate branches of the tree. Everything under 01_Evidence/EV-004/ stays untouched, and anything you make lives in 04_Data/ or 06_Drafts/ with a reference back to EV-004. This is the whole basis of chain of custody: the ability to show that a document you relied on is the same document you received, in the same state.
Where file integrity matters, record a checksum when the material arrives and check it again before you rely on it. Most operating systems will calculate a SHA-256 hash of a file for you, and any reputable checksum tool does the same. A photo of a document, a scanned page, and an emailed PDF should all be hashed on arrival; the log of those hashes is your provenance record.
Also strip or preserve metadata deliberately. A scan carries a device identifier and a timestamp, which is useful provenance and a risk if the file leaves your control. Decide per document which one applies, and write the decision in the register.
How to tell it worked: recalculate the checksums for your originals. Every value matches its arrival record, and nothing under the evidence branch has a modification date later than the collection date.
Build an evidence and source register
The register is the single most useful artifact in the whole system. Folders tell you where a file is; the register tells you what it is, where it came from, and why you kept it. One row per item, one spreadsheet, one owner.
| Field | What goes in it |
|---|---|
| Evidence ID | EV-001, EV-002, sequential and never reused |
| Filename | Exact stored name, including version and status |
| Source | Person, organisation, or public office it came from |
| Provenance | How it was obtained: request, delivery, walk-in, on-site capture |
| Collection date | YYYY-MM-DD the material reached you |
| Document date | YYYY-MM-DD on the document itself, if different |
| Format | PDF, DOCX, CSV, JPG, physical, audio |
| Confidentiality | Public, internal, restricted, source-protected |
| Related claim | The specific assertion in your draft this supports or contradicts |
| Verification | Unchecked, single-source, corroborated, verified, refuted |
| Notes | Gaps, contradictions, follow-up actions |
A filled row looks like this: EV-014, 2026-03-14_INV-2026-014_PortAuthority_Invoice_v01_HOLD.pdf, Port Authority procurement office, public records request PRL-2026-07, collected 2026-03-14, document dated 2026-02-28, PDF, internal, supports the claim that the second tender was issued before the incumbent’s bid was opened, unchecked pending a second copy, note that page 4 is illegible and needs a re-request.
The claim column is what stops the investigation from drifting. Without it you accumulate files; with it you accumulate evidence, and you can see instantly which claims rest on a single unchecked document.
How to tell it worked: run a checksum of the evidence ID column against the filenames. Every file in the evidence branch appears exactly once, and every register row points at a file that exists.
Tag and index the material
Folders, filenames, metadata, and full-text search are four different tools, and the confusion between them causes most filing systems to fail. A folder is a location. A filename is a label. Metadata is structured data stored inside the file. Full-text search reads the contents of the file regardless of where it sits.
Use tags for the things that cut across folders: people, organisations, events, places, topics, and verification status. Tagging is what lets a document belong to a person and a place and a date without being duplicated into three folders. Ten to fifteen controlled tags is plenty, and a short written legend explaining each one prevents tag sprawl.
Run OCR on anything scanned. A searchable PDF can be found by a phrase inside it; a flat image cannot be found at all, which means the document is effectively invisible to anyone who does not already know it exists. Searchable PDFs plus a light folder structure is the combination people in the forums converge on once volume climbs, and it holds up better than deep nested folders.
Where you have text or data, open the CSV in a proper table tool rather than a word processor. It sorts, it filters, and it will not quietly reflow your dates and postal codes into nonsense.
How to tell it worked: search for a distinctive phrase from the middle of a document you filed six weeks ago. If it comes back in under five seconds, indexing is doing its job.
Review, deduplicate, and document gaps
Set a recurring review pass, even if it is one hour a week. Work through the register in order and look for five specific things: missing metadata, duplicate versions, unreadable files, contradictions, and open questions.
Duplicates are usually version problems wearing a disguise. When two files share an evidence ID, compare hashes: if they match, one is a copy and can go; if they differ, keep both and mark the newer one as the working copy with the reason noted. Unreadable files — an audio clip with no transcript, an illegible scan, a broken spreadsheet — get a follow-up action with a name attached, not a silent place in a folder.
Contradictions are the most valuable thing this pass finds. Two documents that disagree, or an interview note that conflicts with a signed record, belong in a chronology with both sides shown and the conflict unresolved. Recording an unresolved conflict honestly is better than quietly dropping the inconvenient document; a reviewer who finds the gap themselves will assume the worst.
Keep a running gaps list with one line per unresolved item and one line per action. “Invoice page 4 illegible, re-request by 2026-04-10” is a complete entry. “Follow up” is not.
How to tell it worked: every register row has either a verification status or a dated follow-up action. There should be no row sitting blank.
Back up and share securely
Use the 3-2-1 approach: three copies of the material, on two different kinds of media, one of them off-site or otherwise disconnected. One working copy on your machine, one synced copy in institutional or organisational storage, and one encrypted archive on external media you do not leave plugged in.
Encrypt the restricted and source-protected branches. Share by named access rather than by open link, so access can be withdrawn when a contract ends or a source asks. Practise least privilege: people get the branch they need for their task, not the whole investigation.
Then run the two tests people skip. Restore one random file from the archive to confirm it opens intact and matches its checksum. Revoke one former collaborator’s access and confirm the revocation took, because shared-drive permissions outlive projects more often than anyone expects.
How to tell it worked: a colleague who has never seen the structure can restore a file and can no longer open a branch they used to have. Both conditions verified, not assumed.
Common Mistakes That Break an Investigation File
Almost every disorganized case file I have seen contains the same handful of failures. Each one has a straightforward fix.
Saving everything to the desktop. The desktop is not a filing system; it is where files land when you have not decided anything yet. Fix: create the parent investigation folder first and drag the scattered material in on day one, even if you have not sorted it yet.
Vague names. A file called “document” or “final” cannot be found, cited, or checked by anyone but you. Fix: apply the date-code-source-type formula retroactively as you touch each file, and accept that renaming takes an hour.
Overwriting originals. Editing the file you received destroys the version you can prove you received. Fix: read-only originals, working copies elsewhere, checksums recorded on arrival.
Relying on folders alone. Deep hierarchies feel organised and search nothing. By week twelve, relevant material is spread across six branches. Fix: searchable PDFs plus tags, with the folder tree kept shallow enough to remember.
Storing sensitive material in personal accounts. A personal drive is outside your institution’s retention, access, and legal-hold rules. Fix: institutional storage for anything with a confidentiality level, and a written access list.
Sharing broad links. An open link to the whole folder cannot be partially revoked and does not tell you who read what. Fix: share named access with the minimum level each person needs, and review the list at handoff.
Never testing restoration. A backup you have never restored is a hope, not a copy. Fix: quarterly, restore one file and compare the checksum.
If you have inherited someone else’s mess rather than building from scratch, work in a fixed order: copy everything into one new parent folder without changing anything, run duplicates and stray files to the top of your review list, build the register from what exists rather than from what you wish existed, and rename only after the register is done. Renaming first loses the trail of which duplicates were actually the same file.
Frequently Asked Questions
What is the best way to organize documents for an investigation?
Use one parent investigation folder, a shallow folder tree built around your reporting question, a naming convention of date plus project code plus source plus document type, and a spreadsheet register logging every document’s source, collection date, and related claim. Search does the heavy lifting, so keep the tree shallow and make the naming consistent. Originals stay untouched in their own branch, with working copies elsewhere.
Should I keep original and edited investigation files together?
No. Keep originals in a read-only evidence branch, untouched, with a checksum recorded on arrival. Put transcriptions, annotations, translations, redacted copies, and analysis extracts in separate working branches that reference the evidence ID of the original. This preserves chain of custody, because you can show the document you relied on is the document you received, in the state you received it.
How should I name files when working with confidential sources?
Use a source code rather than a real name in the filename when the folder is shared more widely than the source’s protection level allows. The mapping from code to name lives only in the protected register, not in the filename. Keep the rest of the convention intact, since date and document type carry most of the search value, and put the sensitive material in a separately encrypted branch.
What metadata should an investigation document register contain?
At minimum: a unique evidence ID, the exact filename, the source, how it was obtained, the collection date, the document date if different, the file format, a confidentiality level, the specific claim it supports or contradicts, a verification status, and notes on gaps or follow-ups. Without the claim column the register is only an inventory. With it, you can see which findings rest on a single unchecked document.
Is cloud storage safe for sensitive investigation documents?
Cloud storage is fine when access is named rather than link-based, encryption is on, the organisation’s retention and legal-hold rules cover it, and access is revoked when the project ends. It is not fine for material you cannot place under institutional control. For the most sensitive sources, keep a separate encrypted archive off the shared platform entirely.
How often should I back up an investigation folder?
Keep three copies on two kinds of media with one off-site or disconnected, and sync working material at least daily while the investigation is active. Test it quarterly by restoring one file and comparing its checksum against the arrival record. A backup you have never restored is a hope rather than a copy, and restoration is the only thing that proves it works.
Start with the register, not the folders. Give the investigation a code, create the parent folder, and log the twenty documents you already have, recording for each one where it came from and which claim it supports. The folder tree and the naming convention fall out of that in about an hour, and from then on how to organize documents for an investigation becomes a habit rather than a project.


