Skip to main content
All articles
Strategy
5 May 2026 4 min read

Designing SEO experiments with keyword clusters: choosing treatment and control groups without contaminating the test

SEO tests can fail when treatment and control pages serve different demand or influence each other. Learn how clusters can improve test grouping while preserving experimental caution.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Two comparable keyword-cluster diagrams labelled Treatment and Control, separated by a boundary that prevents overlapping connections between the groups.

SEO experiments are notoriously difficult because pages do not exist in a vacuum. Unlike a laboratory, your website is a complex web of shared templates, internal links, and overlapping query clusters. External factors like fluctuating search demand and constant algorithm updates further muddy the waters. While keyword clusters can significantly improve how we design these tests by helping us select comparable groups, they are not a magic fix for the inherent messiness of live search data.

The first step is to define exactly what you are testing. Whether you are rolling out new title tag patterns, adding internal link modules, refreshing category copy, or implementing structured data, the intervention needs to be clear. You should ideally test one specific change or a very tightly controlled bundle of changes to ensure the results are actually interpretable.

Using clusters for better matching

To get a reliable signal, your treatment and control pages need to be as similar as possible. Instead of just picking pages at random, you can use clusters to match them based on topic family, page type, historical impressions, and seasonality. This ensures that you are not comparing a high-intent commercial cluster against a low-value informational one, which would make any result meaningless.

A practical example: category titles

Imagine a retailer wants to test a new title tag format across 50 category pages. Using ClusterIQ, they can identify a pool of 100 comparable categories within the same product family that share similar historical demand. Fifty pages receive the new titles (the treatment), while the other fifty remain unchanged (the control). By using cluster data to ensure these groups are truly comparable, the retailer can have much higher confidence that any uplift is due to the title change rather than structural differences between the pages.

Managing query overlap and spillover

One of the biggest risks in SEO testing is contamination. If your treatment and control pages both compete for the same query cluster, an improvement in one might cannibalise the visibility of the other. You should use query-page ownership analysis to ensure your groups are distinct. Similarly, be wary of internal linking. If your treatment involves adding links from control pages to treatment pages, you have fundamentally changed both groups, making the comparison invalid.

Accounting for seasonality and stability

Search volume is rarely flat. Treatment and control groups should ideally exhibit similar seasonal patterns. By looking at cluster-level seasonal profiles rather than just total traffic, you can match groups that react to the market in the same way. It is also vital to check pre-period stability. If your treatment group was already growing at twice the rate of your control group before the test even started, any post-launch success is likely just a continuation of that trend.

Handling Search Console data

When measuring results, Google Search Console data requires careful handling. Because Google aggregates data by page and query using specific rules, you must ensure you are using consistent dimensions and filters throughout the experiment. Whether you are looking at clicks, impressions, or ranking distribution, the way you group the data must remain identical across both the pre-test and post-test periods.

Interpreting the results

It is tempting to see a traffic spike on a chart and claim victory, but causality is hard to prove. A well-matched control group helps, but it does not eliminate every possible confounding factor. When the intervention happens at a cluster level, you should evaluate the outcome at that same level. Look at the total impressions for the topic, the ranking distribution, and how the treated pages' share of the cluster has changed. Moving a single keyword from position 7 to 4 is rarely the whole story.

Maintaining experimental integrity

To keep the process honest, pre-define your success metrics before you see the data. This prevents "p-hacking," where an analyst searches through dozens of metrics until they find one that looks positive. You should also run quality checks to ensure no accidental noindex tags were added or major site migrations occurred during the test window.

In some cases, it makes more sense to use entire clusters as your experimental units rather than individual pages. If a cluster contains several closely related pages, assigning the whole group to either treatment or control can prevent sibling pages from contaminating each other's results.

Building a feedback loop

Every experiment, whether successful or not, is a learning opportunity. By recording results against specific recommendation types in ClusterIQ, such as content consolidation or template changes, you can refine your future strategy. Negative results are particularly valuable; if a theoretically sound cluster opportunity fails to deliver, it should prompt a rethink of your implementation or prioritisation assumptions.

Practitioner principle: clusters can make treatment and control groups more comparable. They cannot remove the need for careful experimental design and modest causal claims.

ClusterIQ Conclusion

Keyword clusters bring structure to the chaos of SEO testing. They allow for better matching, reduced overlap, and more meaningful outcome units. By connecting these experiments back to specific recommendation types, ClusterIQ helps create a continuous feedback loop between your initial analysis and the actual results observed in the SERPs.

Pre-register the cluster selection before looking at the result

Experimental credibility improves when treatment and control selection is fixed before the post-change data is inspected. ClusterIQ can store the eligible page pool, matching variables, exclusion rules, primary outcome and test start date as part of the experiment record. This prevents an analyst from quietly removing awkward pages after launch or switching to a different cluster metric because the original outcome was flat. It also makes repeated experiments comparable because the matching and exclusion logic remains visible.

Where pages cannot be randomised, the limitations should be explicit. A matched comparison can still provide useful evidence, but the result should be described as stronger or weaker observational evidence rather than a perfect causal estimate. The product's role is to improve the design and audit trail, not to make every SEO change look like a controlled laboratory experiment.

Keep implementation dates exact

If treatments roll out over several days, record the actual activation date for each page rather than using one nominal launch date. ClusterIQ can align the observation window with the real implementation so the analysis is not diluted by pages that were still untreated during part of the test.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.