Building a human benchmark set for keyword clustering
You cannot evaluate clustering against vague intuition. Learn how to build a compact human-labelled benchmark of same-group, related-but-separate and unrelated keyword pairs.

Farky Rafiq
Founder of ClusterIQ

You have clustered 1,500 keywords, changed a setting and run them again. The groups look tidier, but are they better for planning pages?
A small, human-labelled benchmark gives you a consistent reference for comparing embeddings, which represent query meaning numerically, grouping thresholds and clustering algorithms. It should reflect decisions that matter in your work.
Start from the downstream decision
For page planning, “semantically similar” is not enough. Similar queries can need different pages. Ask:
Could one page reasonably satisfy both queries under the same page purpose?
For topic discovery, ask instead:
Do these queries belong to the same broader subject area?
Neither question suits every project. Match labels to the job.
Use at least three classes
A binary “same” or “different” label can conceal ambiguity. Use:
- same operational group: queries belong together for your task;
- related but separate: connected, but needing separate groups;
- unrelated: queries do not belong together.
Add “uncertain” when evidence does not support a confident judgement. Reviewers should not have to guess.
Balance hard and easy cases
Random query pairs are mostly obvious non-matches, so deliberately sample:
- paraphrases sharing few words;
- the same topic with different intent;
- changed product model names or numbers;
- geographic modifiers;
- branded versus generic queries;
- ambiguous short queries;
- near-identical wording needing different page types;
- terms connecting otherwise separate topics.
Keep representative clear matches and non-matches too. An edge-case-only benchmark distorts performance and cannot show whether improvements break previously reliable behaviour.
Write annotation guidance
Experienced SEOs may disagree because one considers shared topics while another considers shared landing pages. Write a short guide covering label meanings, search intent, page type, location separation, brands, product models and when to choose “uncertain”.
Guidance will not eliminate disagreement, but helps reveal where the task needs clearer definition.
Measure reviewer agreement
Have multiple reviewers independently label at least a sample, without seeing each other’s answers. Repeated disagreement suggests a definitive model answer may be unrealistic.
Preserve original labels before resolving disagreements, so agreed answers do not hide meaningful uncertainty.
Keep the benchmark outside training and tuning
Repeatedly adjusting thresholds against one small test set risks overfitting: settings improve on those examples without improving elsewhere.
Where practical, separate development examples used for tuning from a hold-out benchmark reserved for final comparison, never training or tuning. Both can be modest; independence matters.
Turn human overrides into tests
Real-work corrections reveal useful distinctions. If marketers repeatedly move “pricing” queries out of informational clusters, add examples of those relationships. The benchmark then records distinctions the product needs to handle.
Evaluate more than one metric
One overall percentage can hide commercially important failures. Track:
- Same-group precision: how often grouped queries genuinely belong together.
- True-paraphrase recall: how many genuine rewordings are correctly grouped.
- Related-but-separate error rate: how often connected queries needing separate groups are mishandled.
- Performance by query type: which kinds cause problems.
- Performance by language or category: whether results hold across markets and subjects.
- Reviewer disagreement rate: how often human judgements differ.
Use regression testing
Run the benchmark after meaningful changes to embeddings, preprocessing such as query cleaning, grouping thresholds, entity-handling rules for brands or products, search intent classification, or edge weighting, the strength assigned to query connections.
This checks for broken behaviour and supports release decisions with evidence rather than appearances.
Keep examples explainable
For every pair, retain the original unprocessed queries, label, reasoning, reviewer confidence, market and relevant page-type notes. Reasoning lets you revisit decisions when future models disagree, rather than treating labels as unexplained rules.
Practitioner principle: include the arguments your SEO team actually has, not just examples the model finds easy.
ClusterIQ Conclusion
A compact human-labelled benchmark makes keyword-clustering comparisons consistent. Build around practical decisions, balance difficult and representative cases, preserve uncertainty and turn corrections into tests. Choose models and thresholds using evidence, not whichever output looks neatest.
How large should the benchmark be?
There is no universal number. A few dozen pairs usually cannot support dependable comparisons; thousands of carefully reviewed examples with disagreements resolved may be unnecessary for an early product.
Cover the main mistakes across important categories, then expand when real-world reviews reveal new errors. Change the benchmark more slowly than the model to preserve meaningful regression checks.
Benchmark whole clusters too
Pairs test similarity and thresholds, but some mistakes emerge only across whole groups. Include manually reviewed clusters recording which terms belong, which do not, and each group’s page purpose or topic.
This tests fragmentation, where a useful group splits apart, and over-merging, where distinct groups combine, beyond pair relationships alone.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

How to evaluate keyword cluster quality: metrics, stability and human review

Cluster stability for SEO: how to tell whether a topic survives small changes
