Check this page with an assistantOpens a chat asking it to summarise this article and name the evidence behind each claim.
Claude opens with the prompt on your clipboard: Anthropic does not support prefilled prompts on the web, and we would rather copy it than ship a button that drops it.
Almost Every Reported SEO Win Is Uncontrolled
The standard claim takes the form "we changed the titles and traffic rose 18%". It contains no control, so it cannot distinguish the change from seasonality, a core update, a competitor's outage, or a demand shift. The number is real; the attribution is decoration.
This matters commercially rather than academically. Uncontrolled results are how teams keep doing things that never worked, and how agencies keep charging for them. If a rewrite coincided with a seasonal upswing, the lesson learned is that rewrites work, and the next twelve months get spent on rewrites. The cost of not testing is not a missing statistic, it is a strategy built on coincidence.
Search makes proper testing harder than it is elsewhere. You cannot serve one version to some visitors and another to others on the same URL, because that is cloaking when the visitor is a crawler. Split testing solves the constraint by changing the unit of assignment: not which user sees what, but which pages get changed.
Every performance claim you make names either a control group or the fact that it has none.
Splitting Pages Instead of Users
Take a set of pages that behave similarly, assign them to two groups, apply the change to one group only, and track organic performance for both. The control group experiences every external condition the variant group does, so divergence between them is attributable to the change in a way that a time comparison never is.
The requirement that does the work is comparability. The two groups need to be similar in traffic volume, page type, template, and seasonality profile before you touch anything. Splitting your top ten pages against your bottom ten is not a test, it is two different populations that would have diverged anyway. Randomised assignment within a homogeneous set is the reliable approach, and stratifying by traffic band before randomising is better still.
This also means split testing is a large-site technique. It needs many pages of the same kind: product pages, location pages, category pages, or a substantial article library. A business with eleven pages cannot run this, and the correct response to that is to say so rather than to run it anyway with eleven pages and report a percentage.
Designing a Test That Can Be Read
Decide everything before you ship: the hypothesis, the metric, the groups, the duration, and what result would make you abandon the idea. Decisions made after seeing data are how tests become justifications.
- State one hypothesis, with a direction.
"Adding a direct answer under the H1 on product FAQ pages will increase clicks", not "improving the pages will help". A hypothesis you cannot be wrong about is not a hypothesis.
- Change exactly one thing.
If you change titles, internal links, and layout together, a positive result tells you that some combination of three things helped. That is not actionable, and it is the most common way tests waste a month.
- Randomise within a homogeneous set.
Same template, similar traffic, similar intent. Order the pages by clicks, pair them off, and assign one of each pair to each group. That gets you balanced groups without needing anything sophisticated.
- Pick the metric that matches the change.
Title changes affect click-through rate, so measure clicks against impressions. Content changes affect ranking, so measure impressions and position. Measuring clicks after a content change conflates two mechanisms.
- Set the duration in advance, then wait.
Two to four weeks after the change has been crawled across the variant group, longer on slow-crawled sites. Confirm the change was actually picked up before you start counting, or your first week is measuring nothing.
- Write down what would falsify it.
Agree beforehand what result means "stop doing this". Without it, every ambiguous outcome becomes "needs more time", and no test ever concludes.
Reading the Result Honestly
You are comparing the change in the variant group against the change in the control group over the same window. If both rose 12%, your change did nothing and the season did the work. If the variant rose 12% while the control rose 2%, you have something worth believing.
| Pattern | Reading | Action |
|---|---|---|
| Variant up, control flat | Plausible effect | Roll out, then re-measure on the next set |
| Both up by similar amounts | External cause, not you | Abandon the change; note what the season did |
| Variant down, control flat | Plausible harm | Revert quickly; this is why you tested |
| Both volatile, no clear gap | Underpowered test | More pages or a longer window, or accept no answer |
| Gap appears then closes | Usually crawl timing, not effect | Extend the window before concluding |
The fourth row is the honest outcome more often than anyone admits. A test that cannot detect an effect has told you the effect is not large relative to your noise, which is genuinely useful information and is not the same as the change being worthless. Report it as "no detectable effect at this sample size" rather than as a failure or, worse, as a win by squinting.
More sophisticated approaches exist and are worth knowing about without being required. Practitioners running these at scale use time-series models that learn the relationship between control and variant before the change and then forecast what the variant would have done without it. That is better than eyeballing two lines, and eyeballing two lines is dramatically better than no control at all.
When You Cannot Split, Say So
Most small businesses genuinely cannot run a split test. Too few comparable pages, too little traffic, too much week-to-week variation. The correct response is to be honest that changes are unmeasured rather than to run an underpowered test and dress the output up.
There is still a discipline available. Change one thing at a time, record the date, and wait long enough to see whether anything moved before changing the next thing. That is not a controlled experiment and it will not survive a seasonal swing, but it produces a timeline you can reason about, which is the minimum needed to diagnose a decline later.
The other option is to test where you can and reason where you cannot. Title and description changes can sometimes be tested on a large enough article set even when your commercial pages are few, and the findings frequently transfer. What does not transfer is a claim from someone else's case study on a different site in a different market, which is worth remembering every time a tactic arrives with a percentage attached.
The Change Log Is the Cheapest Half
One dated row per change, naming what changed, which URLs, and who did it. Almost nobody keeps one, and its absence is why most traffic investigations end in speculation about algorithm updates.
The value shows up months later. When traffic drops, the first question is what changed and when, and a log answers it in a minute. Without one you are reconstructing history from memory and deploy timestamps, and the temptation is to blame whatever update happened nearby, which is how sites end up chasing an algorithm when they shipped a noindex.
Keep it beside the rest of your search reporting rather than in a separate document nobody opens. It costs a line per change, it makes every future diagnosis faster, and it is the prerequisite for the decline investigation described in page-level authority and core update recovery. The reporting structure it belongs in is in the SEO reporting framework.
Questions People Ask About SEO Testing
- What is SEO split testing?
It is A/B testing adapted to a constraint: you cannot show different versions of a page to Googlebot without cloaking, so instead of splitting users you split pages. Take a set of comparable pages, apply a change to half of them, leave the other half untouched, and compare how organic performance diverges between the two groups over time.
- Why can't I just compare before and after?
Search traffic moves for reasons unrelated to your change: seasonality, algorithm updates, competitor moves, and demand shifts. Before-and-after comparisons attribute all of that movement to your edit. Control groups experience the same external conditions, so the difference between groups isolates your change.
- How many pages do I need to run a test?
Practically, at least 50 to 100 pages split into two groups, with a couple of hundred giving you a much better chance of detecting a modest effect. Below roughly 50 the natural week-to-week variation between page groups will usually swamp any change you make, and you will get a result that looks decisive and is not.
- How long should an SEO test run?
Two to four weeks after the change is crawled and processed is a common range, and slower-crawled sites need longer. Two things spoil short tests: the change may not have been picked up on every page yet, and you may be reading a week that was unusual for unrelated reasons. Decide the duration before you start and resist stopping early on a good-looking trend.
- Is SEO split testing cloaking?
No, as long as every page serves the same content to everyone. Cloaking is showing different content to crawlers than to people on the same URL. Page-group splitting changes some pages for all visitors and leaves others alone, which is ordinary site development with a measurement plan attached.

