To run an A/B test on headlines, you serve two or more headline versions of the same story to randomly split readers, keep everything else identical, and measure which version produced the outcome you actually care about. Eight steps cover it: define the decision, pick the metric, write the variants, size the audience, randomize and preserve the assignment, instrument the events, launch without peeking, and read the result.
Most newsrooms already have the raw material for this. You publish dozens of stories a day, your analytics already know how often each headline on the homepage gets clicked, and you have two or three drafts sitting in the CMS that you argued about in Slack. What is missing is usually not tooling but a rule: one question per test, one primary metric, and a stopping point you commit to before the data arrives.
This guide is written for journalists, editors and audience teams at digital publishers, and for independent bloggers and newsletter writers doing the same job with a smaller team. It treats headline testing as an editorial workflow rather than a marketing exercise, which changes what you measure and what counts as a win.
Table of Contents
- What You Need
- Step-by-Step: How to Run an A/B Test on Headlines
- 1. Define the decision the test should make
- 2. Choose the metric before choosing the variants
- 3. Write genuinely different headline variants
- 4. Calculate the audience and test duration
- 5. Randomly assign readers and preserve the assignment
- 6. Instrument events and validate tracking
- 7. Launch the test and avoid early decisions
- 8. Read the results and make an editorial decision
- Common Mistakes
- Frequently Asked Questions
- Conclusion
What You Need
Six things have to be true before you send a single reader to a variant.
- A hypothesis in one sentence. If more than one headline idea would satisfy it, split it into separate tests.
- At least two genuinely different variants. A punctuation change is not a variant, it is a typo test.
- An eligible, comparable audience. You need to know who is in the test and who is excluded.
- Consistent distribution. Same placement, same promotion, same time of day for both arms.
- Consent-compliant analytics. Your click and engagement events need to survive your consent setup without losing half the sample.
- A mechanism that assigns readers randomly. This is the part most hand-rolled setups get wrong.
That last item is where platforms differ most. Newsroom analytics suites built for publishers tend to handle it natively, because they already know which story occupies which homepage slot. smartocto’s Smartify module, for example, surfaces the handful of articles most worth testing right now and opens a headline test from that list, so the editor never has to pick a story by instinct. Google Analytics 4 can run experiments on a properly tagged property, though you build the headline variants yourself and the split applies to sessions rather than to an editorial slot. Most CMS platforms have no randomizing feature at all unless an experimentation module is installed.
Pick based on where the randomization has to happen. If the decision is which headline to run on the homepage right now, the test belongs in a tool that can control a homepage slot and expose click and engagement data for that slot. If the decision is about a newsletter subject line or a landing page, an experimentation tool with conversion tracking fits better. Whichever you choose, the newsroom still owns the hypothesis, the variants and the accuracy check. No tool can rescue a headline that overstates the story.
Two constraints deserve a decision up front. If your analytics consent banner blocks a large share of sessions before events fire, you are effectively testing only the readers who already consented, which skews toward loyal visitors. And if a headline test would swap the headline on a URL that has already been indexed and shared, treat that as an editorial and search question, not a silent optimization. Decide who signs off on it, and keep a record of the original and the promoted version.
Step-by-Step: How to Run an A/B Test on Headlines
The eight steps below follow the order that actually prevents bad decisions. Most headline tests that go wrong skipped step one or two, then tried to rescue the result in step eight.
1. Define the decision the test should make
Write the decision first: “we will use whichever headline framing produces more engaged reading time on this story.” That sentence is your whole brief.
A test that only asks “which headline gets more clicks” will find the most sensational version every time, which is not the same as finding the best headline. Publisher analytics vendors have long argued that headline effectiveness should be read across click-through rate, average time on page and recirculation rate together, because a headline can win the click and lose the reader in the same breath.
Write down what the winner will inform before you start. If the answer is a different framing for the same story, keep the finding. If the answer would rewrite your headline standards, that is a much bigger claim than two days of homepage data supports.
2. Choose the metric before choosing the variants
Pick one primary metric and one or two guardrails, then write them down. The primary metric is the thing that decides the test. The guardrails are the things that can veto it.
The usual options for a news site are qualified click-through rate on the homepage slot, average engaged time on the story, recirculation rate, subscription conversion, and for a newsletter, open and click rates. If you sell subscriptions, conversion is a reasonable primary metric. If you are building an audience habit, engagement time and recirculation matter more than raw clicks, because a reader who clicks once and leaves has not become a regular.
Guardrails catch the classic failure mode. A curiosity gap headline can lift clicks by a wide margin while average time on page drops, which usually means the headline promised something the story does not deliver. Track bounce rate or negative feedback as a veto signal, and decide in advance what magnitude of regression makes you reject a winner.
One primary metric only. Two co-primary metrics give you permission to declare victory on whichever one flatters the story, which is how a test ends with a conclusion nobody planned for.
3. Write genuinely different headline variants

Variants should differ on the editorial variable you are testing and agree on everything else. Same story, same photo, same slot, same promotion. If two versions differ in wording style and length at once, you will learn nothing when one wins.
Pick the variable deliberately. Some pairs worth testing on a news site:
- Specificity against vagueness. “Council approves 40-home housing plan” against “Big housing plan approved”.
- Benefit or consequence framing. What the reader learns or what changes for them.
- Direct against curiosity gap. The second is allowed only if the story genuinely delivers the withheld detail.
- Number against no number. Lists and counts often help scanners on a crowded homepage.
- Geographic framing. Adding a place or demonym is a cheap, well-documented way to test local relevance against general appeal.
Research on the Upworthy headline archive, built on thousands of paired headlines the publisher ran as real tests, found no single stylistic feature reliably predicted success. That is the useful takeaway: variant generation is a discipline of listing framings, not a hunt for a magic word.
Before launch, run an accuracy check on every version. Does it describe what the article actually says, does it avoid a claim the story does not support, and would you defend it to the person at the center of the story? A headline that fails is disqualified regardless of how it performs. Ambiguity is the same problem in reverse: if two readers would walk away with a different understanding of the story, the headline is not doing its job, however well it tests.
4. Calculate the audience and test duration
Decide how many readers each variant needs before you look at any results. Otherwise you will stop at whatever point the difference looks convincing, which is exactly how false winners get promoted.
Three inputs determine the number: your current baseline rate, the smallest difference you would actually act on, and the confidence you require. The smallest difference matters most, and it is the one people skip. Deciding in advance that you will only act on a change of at least one percentage point stops you from chasing a 0.2 point wobble that will reverse next week. With a baseline and a minimum detectable effect fixed, an A/B test sample size calculator will give you the required impressions per variant in seconds. You do not need to understand the mathematics to use it; you do need to fill in honest inputs.
Then set a minimum runtime that spans a full week, because weekday and weekend traffic behave differently on a news site. A test that runs Monday to Thursday mostly measures the working week. If your publication cannot reach the required sample in a week, do not shorten the test below the minimum. Extend the runtime instead, or change what you are testing.
Low traffic is a real limit, not a technicality. One publisher tool’s guidance suggests headline testing pays off at meaningful scale, on the order of a hundred thousand homepage views a day, which puts it out of reach for most independent publications and many regional ones. If you are below that, your options are to test on your highest-traffic story only, to test one story at a time on a homepage slot, to accept a much longer runtime, or to skip the test and rely on a documented editorial process. Running an underpowered test and treating the result as evidence is the worst of the four.
Set the rule for sequential monitoring separately. If your platform supports continuous monitoring with an option to stop early, you can define an interim rule with a stricter threshold. Keep it fixed in advance and do not adjust it once you have seen partial data.
5. Randomly assign readers and preserve the assignment
Split eligible traffic 50/50, or use another allocation only when you can justify it. Randomization has to happen before the variant is rendered, and the same reader must keep the same variant for the whole session, ideally across sessions.
That last point matters more than it sounds. If a reader sees the control headline on the homepage and the variant inside the article, your click event and your engagement event belong to different experiences, and the comparison collapses. Assign once, store the assignment against a stable identifier, and carry it through the click to the landing page and the conversion event. Most experimentation tools handle this with a cookie or an equivalent identifier; if you are hand-rolling it, that persistence is the first thing to build.
Worked example: your homepage has 6,000 eligible sessions a day. A 50/50 split gives 3,000 per arm, so a test requiring 20,000 impressions per variant needs about seven days. If the split is skewed toward the control because the assignment logic falls back to the default when the identifier is missing, the control silently inflates and every result is suspect, so check the observed split rather than the configured one.
When your CMS cannot randomize, the fallback is weak but workable. Rotate two headline versions across homepage slots by a fixed schedule over alternating days, keep the slot position and the promotion time constant, exclude direct traffic that arrived through another story’s link, and analyze it as a day-alternating comparison rather than a true random split. Say plainly in your write-up that it is not a clean experiment, and treat the result as directional.
6. Instrument events and validate tracking
Map four events before launch: headline impression, headline click, qualified engagement on the story, and the conversion you care about if you have one. Confirm that both variants record the same event definitions, with the same parameters and the same exclusions.
The variant identifier has to travel with every event. If you only tag clicks and not impressions, you cannot compute a rate, you only have raw counts. Impressions are the denominator, and on a homepage slot the temptation is to count a view of the whole page as a headline impression, which inflates one variant if one version renders taller than the other.
Platform-neutral naming is fine as long as it is consistent. Most analytics systems expose something like a page view, an outbound click or a scroll event, and consent management platforms layer their own labels on top, so the exact menu path differs between Google Analytics 4, your CMP and your CMS analytics screen. What matters is that your test report can state, for each event, where it was defined and who confirmed it.
Then validate before launch. Load the page as a logged-in editor and as a logged-out reader, confirm both variants render, trigger the click and the engagement event manually in each, and check the numbers arrive attributed to the right variant. Exclude internal traffic from your own office IP range, filter obvious bots, and deduplicate. A quality-assurance pass that takes an hour routinely catches a misfiring tag that would otherwise have run for a week.
7. Launch the test and avoid early decisions
In the first 24 hours, check rendering, the traffic split and event arrival, not the outcome. Then stop looking at the performance numbers.
Pre-commit to your stopping rule: either a fixed duration with a sample threshold, or a platform’s sequential test with a predefined rule. What you must not do is change the rule after seeing partial results, because the reader who knows where the data was leaning will keep watching until it crosses a line, and that is precisely the behavior that produces false positives at a much higher rate than the nominal error rate suggests.
Pause for an incident, not for a result. If one variant renders broken on mobile, if a promotion sends a burst of traffic to only one arm, or if an editor accidentally edits a variant mid-test, pause, fix the cause and restart the clock. Restarting costs days. Publishing a headline that was never fairly tested costs credibility, and on a news site credibility is the whole product.
8. Read the results and make an editorial decision

Compare the primary metric and every guardrail by variant, and report both the absolute and the relative difference. An absolute change of 0.4 percentage points on a rate that moved from 4.1 to 4.5 is a ten percent relative lift, and the first number is the one a homepage editor can actually act on.
A simple results record looks like this:
Variant A: 12,400 impressions, 4.1% click-through rate, 38 seconds average engaged time, 11% recirculation
Variant B: 12,380 impressions, 4.6% click-through rate, 34 seconds average engaged time, 10% recirculation
Decision rule: promote B only if the click-through difference clears the threshold and average engaged time falls by no more than 5%
Run that example honestly. Variant B wins on clicks and loses on engaged time and recirculation, which is the pattern where the headline overpromised. Under the rule written in advance, the test is a failure and the control stays. That is a real result and it is worth as much as a win, because it tells the desk something about the framing.
Then classify the outcome honestly into one of three states: a winner that clears your threshold on the primary metric without breaching a guardrail, a tie where the data cannot separate the variants, or a loss where the variant is worse. A tie means keep the control and either stop or run a different test. Declaring a winner from an inconclusive result is the most common way a headline testing routine loses credibility internally, because the next test’s result will not match the confidence everyone now expects.
Write the log entry the same day: hypothesis, metric, guardrails, variant list, sample, result, decision, and who decided. Six lines. That log is the only way the findings accumulate into a reusable pattern instead of a one-off curiosity, and it is what lets you answer whether a framing has worked before.
Common Mistakes
Most failed headline tests come down to a handful of repeat errors.
- Testing several things at once. Different length and a different framing in one pair, and a win tells you nothing usable. Fix: one variable per test, and if you have several hypotheses, queue them as sequential tests.
- Editing a variant while the test runs. Fix: lock the variants, and if one must change for accuracy, invalidate the test and restart.
- Stopping as soon as the numbers look good. Fix: pre-commit to a duration and a sample threshold, and treat early stopping as an exception that needs a written reason.
- Comparing rates built on different sample sizes. A 12% click-through rate on 300 impressions is not comparable to 11% on 30,000. Fix: report the sample next to every rate, and refuse to compare across unequal exposure.
- Optimizing for clicks alone. Fix: always pair the primary metric with an engagement or conversion guardrail, and write the veto rule before launch.
- Letting placement bias the result. Two variants shown in different slots, at different times, or to different channels are not a test. Fix: hold slot, time and promotion constant, and verify the observed split.
- Testing the same audience repeatedly without a record. Over time you learn about the tool’s quirks rather than about headlines. Fix: keep the log, note repeat exposures, and vary the stories.
- Declaring a winner from an inconclusive result. Fix: three outcomes, not two, and a tie is a legitimate result with its own next action.
A few governance habits go a long way. Assign one owner per test so the hypothesis, the decision and the log entry stay with one person. Review the log monthly, because a pattern across a quarter of tests is far more useful than any single result. And keep the accuracy standard outside the test entirely: the winning headline still has to be true, and no engagement number should be allowed to argue otherwise.
One more caution worth raising with editors. Swapping the headline on a URL that has already been published and indexed is not purely a testing decision. It changes what search engines and social platforms may have already stored, it can make an incoming link point at a page that no longer looks like the one it linked to, and it invalidates the shared preview card. Some publications simply do not run headline tests on stories that are already widely distributed, and instead test on a story’s first hours of homepage life.
Frequently Asked Questions
How long should an A/B test on headlines run?
Run a headline test for at least one full week so weekday and weekend traffic are both represented, and never end it before the sample size your baseline and minimum detectable effect require. If the sample cannot be reached in a week, extend the runtime rather than shortening it. Fixed-duration tests are simplest; sequential monitoring platforms let you set a stopping rule in advance instead, but the rule must not change once results start arriving.
Should the same reader always see the same headline variant?
Yes. Assign a reader to one variant once and keep that assignment for the session, ideally across sessions using a stable identifier. If someone sees the control on the homepage and the variant inside the article, their click event and their engagement event describe different experiences and the comparison breaks. Most experimentation tools handle persistence automatically, which is one reason a real platform beats hand-rolled randomization.
Is click-through rate enough to choose a winning headline?
Not on its own. Click-through rate measures attention, not satisfaction, and the highest-CTR headline often pulls short visits and a higher bounce rate. Pair it with average engaged time, recirculation or subscription conversion as a guardrail, and write the veto rule before launch. On a news site the headline that wins clicks while losing engaged time is usually promising more than the story delivers.
When is a headline test result statistically significant?
A result is significant when the chance of seeing a difference that large if there were no real difference is below your chosen threshold, usually five percent, which is the same as saying you are at least 95% confident the difference is real rather than noise. Significance requires enough sample on both variants to detect the minimum difference you care about, which is why sample size and runtime must be fixed before launch.
What if our CMS does not have built-in A/B testing?
You have three options: install an experimentation module that adds randomization and event tracking, move the test into an analytics platform that supports experiments while your CMS serves both variants, or run a documented fallback such as day-alternating rotation of two headlines in the same homepage slot. The fallback is weaker because it is not randomly assigned, so report it as directional and never quote it as a clean experiment.
Conclusion
The first thing to do is write one sentence: the decision this test will inform. Then pick one primary metric with a guardrail beside it, write two headline variants that differ on exactly one editorial variable, and commit to a sample size and a stopping rule before anything goes live.
Everything after that is discipline. One variable, random assignment held steady through to the conversion event, identical placement and promotion, no peeking, and a log entry on the day of the decision. Run it that way and each test leaves you with a defensible answer about framing, which is worth far more to a newsroom than a clever headline that happened to spike for two days.


