A usability test on a news site is a moderated session where real readers complete realistic news tasks on your homepage, article template, topic hub or paywall while you watch where they hesitate, backtrack and give up. The whole thing runs about two weeks of prep plus five to eight 45-minute sessions, and it costs little more than the thank-you you give participants.
Analytics tell you a reader left the page after nine seconds. They never tell you why. That gap is the entire reason to run a usability test, and it is the reason publishers with no research team can still do it: you need a moderator, a task list and five people who actually read the news.
There is no news-industry result ranking in the top ten for this query. Every guide out there is written for SaaS onboarding or ecommerce checkout, which is why so many of them miss the things that actually break on a news site: a cookie banner that covers the headline, a paywall that appears mid-paragraph, a live blog with no “latest update” time, a related-story module that keeps sending people back to the same homepage slot.
What follows is the protocol I would hand a product editor or a developer on their first round, in the order they would actually do it.
Table of Contents
- What You Need
- Step-by-Step
- Common Mistakes
- Frequently Asked Questions
- How many people do I need to run a usability test on a news site?
- Can I test a live news site, or do I need a prototype?
- How long should each usability testing session last?
- What should I do when a participant cannot complete a task?
- Should I give usability-test participants instructions or hints?
- Does usability testing replace analytics and reader feedback?
What You Need
Six things, and none of them expensive. Work through this list before you book anyone, because a missing item here quietly ruins the sessions later.
- One specific thing to test. A page, a template or a flow. “The whole site” is not a test; “the redesigned article template on mobile” is.
- Three to five task scenarios written as reader goals, not as interface instructions.
- Participant criteria — who counts as a real reader for this test, including device and reading frequency.
- A recruiting route: your newsletter list, social followers, a panel, or readers who answered a past reader survey.
- A recording and note setup: screen share plus a second person taking notes, or a recorder plus a single observer sheet.
- Consent and privacy wording covering recording, data retention and withdrawal.
- Somewhere to put findings: a spreadsheet with one row per observation, columns for task, severity and evidence.
Budget note: the only real cost line is participant incentives. Researchers in the r/UXResearch community treat a paid thank-you as standard, not optional, and they are right — unpaid news readers are disproportionately the people who already have strong opinions about your site.
Step-by-Step
Each step below should produce something you can point at: a written question, a task list, a filled screener, a printed script, a recording, a severity table, a report. If a step produces nothing, skip it rather than running the session anyway.
1. Define What You Want to Test
Start by turning a complaint into a question with four parts: the page, the audience, the task, and the decision you will make with the answer. “People say the new layout is confusing” fails all four. “Can a reader who visits three times a week find the latest update to a story they read yesterday on a phone, without opening the menu, on the new article template?” passes all four.
The fourth part is where most newsroom tests go wrong. If no decision is attached, the research gets politely filed and nothing changes. Write down the decision first: do we ship this template, do we move the newsletter module, do we keep the metered paywall as is.
Keep usability separate from editorial and visual preference. Whether the headline font is elegant is not a usability finding. Whether a reader can tell which story is the newest one is. Say so explicitly in the test plan, because otherwise the loudest opinion in the room wins.
2. Choose Realistic Reader Tasks
Write three to five tasks that map to goals readers already have, and give each one a success criterion you can judge objectively. Three to five is the working range: fewer and you miss whole journeys, more and attention drops and the later tasks get shallow answers.
A news-site task library looks like this, with the wording you can copy:
- “You want to know what happened in the last few hours. Find the most recent update.” Success: reaches a live or continuously updated story and can state its latest update time.
- “You read a story yesterday and want to see what has been added since.” Success: opens the same story and locates the update marker or timestamp.
- “You want to follow this topic without following the whole site.” Success: reaches a topic hub or tag page from within the article.
- “You have decided to start a subscription. Find out what it costs and what you get.” Success: reaches pricing or signup details, or explains clearly why it is not available.
- “You were interrupted. Come back to the story and find your place.” Success: returns to the story and finds the point they stopped at, or the headline context they need.
Say what the reader wants, never where to click. Any sentence containing “click the menu” has already done the participant’s job for them, and you will get a clean result that tells you nothing about whether the interface works.
3. Recruit the Right Participants
Recruit people who read news in the way the reader you care about reads news. If the test is about the mobile article template for occasional readers, screen for frequency of reading and for phone ownership. If it is about the newsletter funnel, screen for people who already subscribe to at least one publication.
Screen with neutral questions. Ask “How often do you read news online?” not “Do you use our site?” A screener that mentions your brand attracts fans and critics, and both behave differently from a real reader. Adding a few plausible wrong options is a standard screening technique for filtering out people who just want to be helpful.
Test with colleagues and friends only as a dry run. Researchers on r/UXResearch repeatedly warn about user proxies: insiders know where the search box is, they read the way you wrote it, and their findings flatter the design. Real problems disappear the moment someone who works there stops helping.
Five to eight participants is a legitimate first round. Nielsen’s five-user heuristic says most usability problems surface in the first five people, and Macefield’s 2009 participant-count research gives a more careful curve behind that shorthand. Treat five as the point where repeated problems emerge, not as a magic number that ends your research.
4. Prepare the Test and Moderator Guide
Fix the environment so differences mean something. Decide device, browser and viewport once and write it in the test plan, and check that your test pages exist in that exact state, including the consent banner, ads and paywall. A cookie dialog that only appears for new visitors will ruin the first task for everyone if you are not careful about it.
Have someone else run the sessions if you can. The moderator watches the screen and talks to the participant; a second person takes notes. Trying to do both is the single most common reason note quality collapses.
Write a one-page script with the neutral prompts in advance. You need wording for the opening, for handing over a task, for silence, and for closing. Keep the whole run-of-show to about 45 minutes: five for the intro, five per task, ten at the end for questions and a satisfaction questionnaire.
Consent wording should be read out and, if you are recording, shown on screen before the recording starts. Say what you are recording, why, how long you keep it, and that the participant can stop at any point. Do not start capturing until they agree.
5. Run the Session and Watch for Behavior
Introduce the session in two minutes: you want to hear what they think, there are no wrong answers, and you are testing the site, not them. Then ask them to think aloud and stop coaching. The moment you explain where the link is, you have lost the finding.
Log what you see, not what they say they would do. Hesitation, backtracking, scrolling without reading, repeated clicking in the same region, and a long pause before an action are all evidence. So is the recovery — a reader who finds it after three wrong turns still had a problem, it just was not a blocker.
Separate difficulty from preference constantly. “I would never use the search bar” is a preference. Two participants scanning the header for three minutes before finding search is a difficulty. Only the second one goes in the findings table.
Budget real time for review afterwards. Takt’s guide puts it at roughly an hour of video review per hour of session recording, which matches what working sessions have always cost. Schedule it before you book the next participant, not after the last one.
6. Analyze What Happened
Do this within a day or two, while the sessions are still sharp. One row per observation: participant, task, what happened, how many people hit it, what it cost them in time, and a severity rating.
Keep the measures simple enough that you actually compute them: task completion rate out of five, time on task for those who finished, and first-click accuracy on the entry point. That is enough to show a pattern, and a pattern that four of five people hit is a far stronger argument to an editor than any single session quote.
Rate severity on the familiar three-tier scale. Critical means the task could not be completed. Serious means completion took wrong turns or a large detour. Minor means it worked but produced hesitation or doubt. Then group observations by cause, not by session, so the report shows that four different people tripped on the same thing.
Bring analytics into the same conversation. If your recirculation rate is low and four participants could not tell a related story from an advertisement, you have a pattern and a number pointing at the same problem. That pairing is what moves a finding from “interesting” to “fund this”.
7. Prioritize and Report the Findings
Rank issues by reach first, then by cost to fix. A small template change that blocks a core task on mobile outranks a beautiful problem with the footer that affects nobody. Write the list so a developer can see scope and an editor can see newsroom impact without decoding anything.
Write each problem as a sentence with evidence attached, not a bullet of adjectives. “Four of five participants could not reach the live blog from the homepage without opening the menu, averaging 38 seconds and two wrong turns; our homepage-to-live-blog click-through is below the article-to-live-blog rate, which is consistent with a navigation problem rather than a demand problem.” That is a paragraph an editor can act on, because it names the cost.
Say who owns each fix and what it depends on. Label items as product, design, editorial or audience, since a finding about a topic hub is often an audience desk decision rather than a design change. Then schedule the retest before the fixes land, not after, so the follow-up round is already booked.
If you want the small version of all of this, a single remote round of five sessions, a written plan and a report gets you most of the value of a full program. The gap between publishers who seem to read their readers and those who do is usually a written protocol.
Common Mistakes
Most failed news-site tests fail in the same six ways, and every one of them has a cheap correction.
Testing without defined tasks. If you ask participants to “look around the site”, you get opinions about colour and nothing you can act on. Correction: three to five tasks with written success criteria, every time.
Recruiting only insiders. Colleagues, friends and the author of the feature all navigate the way they built it. Correction: screen externally first, and treat any internal run as a rehearsal.
Leading the participant. Saying “it’s the blue link on the right” mid-task ends the session for that task. Correction: neutral prompts only — “what would you try next?” and then silence.
Changing the page during the test. If you publish a template fix after session three, sessions one to three are not comparable to four and five. Correction: freeze the build, or run each round in a single day and note the time of every change.
Treating every complaint as a priority. A report with thirty findings has no priorities in it. Correction: severity plus reach, and a short list a team can finish in a sprint.
Declaring success from one session. One clean run is an anecdote. Correction: repeat a problem across three to five people before it enters the report as a finding.
A few things that make the rounds go better: test the metered paywall with people who have not hit the meter that day, since the second and third view behave differently from the first. And never test during a genuine breaking-news moment if you can avoid it, because a live homepage under real urgency is a different product and your participants will tell you about the news instead of the site.
One more, specific to publishers: decide in advance whether you are testing the current site or a prototype, and be honest about the fidelity. People forgive a rough prototype for looking rough. They do not forgive a broken one for asking them to finish a task.
Frequently Asked Questions
How many people do I need to run a usability test on a news site?
Five to eight participants is a workable first round for a news site. Nielsen’s five-user heuristic holds that most usability problems surface within the first five people, and Macefield’s 2009 research gives a more careful version of that curve. Five will find the obvious blockers; eight helps you spot whether a problem is isolated or repeats. If you are testing two distinct reader types, run five per group rather than ten in one mixed session.
Can I test a live news site, or do I need a prototype?
Test the live site whenever the change is small, a layout fix, a paywall tweak, a new module. Readers respond to real headlines, real ad load and real speed, and none of that survives a prototype. Use a prototype only for large redesigns where the build is not ready or would be expensive to change mid-round. Whichever you choose, freeze the version for the duration and write down the date and time of every change.
How long should each usability testing session last?
Plan for 45 minutes with a 60-minute slot booked, so late starts and extra questions do not cut the last task short. Roughly five minutes for the intro and consent, five minutes per task across four or five tasks, and ten minutes at the end for open questions and a short satisfaction questionnaire. Longer sessions produce tired participants and shallower answers on the tasks that run last.
What should I do when a participant cannot complete a task?
Let it run out. Note the point where they gave up, what they tried first, and what they said, then move to the next task. Do not offer a hint and do not show them the answer, because both destroy the evidence. Afterwards, mark the task as failed, note the number of participants who hit the same wall, and rate it critical. A failure that four of five people hit is the strongest finding in the whole report.
Should I give usability-test participants instructions or hints?
No instructions, and no hints. Instructions teach the interface, which is the opposite of what you are there to learn, and a hint ends the evidence for that task. Give a short neutral prompt such as what would you try next, or is there anything else you would look at, and then stay quiet. If a participant asks you for help, acknowledge the request, tell them you cannot help with that task, and move on to the next one.
Does usability testing replace analytics and reader feedback?
No, and pairing them is where the real value sits. Analytics show what happened at scale, such as low recirculation or a 30% mobile drop on one template, but never why. Usability sessions show the why in detail, across only a handful of people. Reader emails and support tickets tell you which issues are already annoying enough to write about. Read the three together and the finding carries a number, a mechanism and a frequency.
Start with one page, three tasks and five readers. Write the test question and the decision attached to it on a single sheet, recruit outside your own building, and run the sessions inside a single week so the site stays still while you watch.
Then take the one problem four people hit and write it up with a number attached. That single paragraph does more for a newsroom redesign than another month of redesigns will, and it is the part of how to run a usability test on a news site that most teams skip.
Redesigns will keep shipping after 2026, and the ones with a written protocol behind them keep getting less broken.


