Skip to main content
All articles
Clustering
24 September 2026 6 min read

How to evaluate keyword cluster quality: metrics, stability and human review

A clean-looking cluster is not necessarily useful. Learn how to combine internal metrics, stability testing and practitioner review to evaluate keyword clustering properly.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing a keyword cluster with a stable dense core, shifting boundary nodes and overlapping alternate cluster outlines, assessed through metrics, stability and task fit.

You have grouped 1,500 keywords into tidy clusters. The labels look sensible, and the chart looks convincing. But can you use those groups to decide which pages to create, update or combine?

A keyword cluster can look good and still be unhelpful. Closely packed points on a chart, a high silhouette score or clear topic labels describe parts of the result. None proves that the groups support the SEO decision you need to make.

To judge cluster quality properly, combine mathematical checks, tests of how much the groups change and practitioner review against the job you want them to do.

Start by defining what “good” means

Before choosing a metric, decide what you need the clusters for.

You might want to:

  • remove duplicates or standardise wording in a keyword export;
  • explore groups of keywords with related meanings;
  • identify possible targets for individual pages;
  • build a content taxonomy, or a structure for organising topics;
  • map queries to existing URLs;
  • find gaps or overlaps in your content programme.

The same grouping can work well for one job and badly for another. A broad group about accounting software might be useful for discovering topics, but too broad to tell you which URL should target each query.

This is why our keyword clustering versus topic clustering distinction starts with the decision rather than the algorithm.

What internal clustering metrics can tell you

Internal metrics check how a clustering result is arranged: how close its members are, how separate its groups are or how they connect. They do this without needing an external set of labels that says which grouping is correct.

One common example is the silhouette coefficient. Scikit-learn calculates this for each item using its average distance to other members of its own cluster and its average distance to the nearest other cluster. Scores range from -1 to 1. Higher values generally suggest clearer separation, values near zero suggest overlapping boundaries, and negative values can indicate that an item is closer to another cluster than to its assigned one.

That makes silhouette a useful diagnostic. It can help you compare settings when you keep the keyword representation and distance assumptions the same.

What it cannot tell you is whether those groups match search intent, suit a particular page type or have commercial value.

Why a good silhouette score can still produce bad SEO groups

Imagine an embedding model, which represents text as numerical vectors, places all queries about accounting software close together. A clustering algorithm then separates that area neatly from payroll, invoicing and banking topics.

The accounting software cluster might be tightly grouped while containing:

  • “best accounting software”;
  • “accounting software pricing”;
  • “free accounting software trial”;
  • “how does accounting software work”.

These queries clearly share a subject. That does not necessarily make them a good target for one page.

Internal metrics assess the structure produced by the keyword representation. They do not know whether a searcher expects a comparison, a pricing page, a trial sign-up or an explanation.

Graph metrics have the same limitation

Graph-based clustering treats keywords as nodes and the relationships between them as connections, or edges. Metrics such as modularity and community connectivity can help assess how well the groups fit that network.

But the network already reflects choices you have made. Which keywords become nodes? How similar must two queries be before you connect them? How much weight does each connection carry, and which similarity signals create it?

A well-separated set of communities in a poorly built graph can still misrepresent your keywords. Review the graph-construction rules as well as the groups they produce.

Stability is one of the most useful tests

A cluster is more convincing when its core members stay together after reasonable changes to the process.

Useful changes to test include:

  • slightly higher or lower similarity thresholds;
  • different random seeds, which affect random choices within the algorithm;
  • small changes to HDBSCAN clustering parameters;
  • nearby graph resolution settings, which influence how broadly or finely communities are divided;
  • alternative embedding models;
  • subsamples of the keyword set.

You do not need every query to stay in exactly the same group. Queries near a boundary are likely to move.

What matters is whether the main topic structure remains recognisable. If a tiny setting change reorganises almost everything, the result deserves a closer look.

Measure movement at the right level

Comparing cluster IDs between runs will not tell you much, because those labels are arbitrary. “Cluster 4” in one run could contain the same queries as “Cluster 11” in another.

Instead, compare which keywords stay together. Ask:

  • Which query pairs repeatedly remain in the same group?
  • Which core topics persist?
  • Which queries often switch communities?
  • Which groups repeatedly split or merge?
  • Which queries are consistently treated as noise, rather than assigned to a cluster?

Unstable queries can be particularly useful to investigate. They may have ambiguous intent, connect two subjects or expose weak assumptions in the modelling.

Create a small human judgement set

If you use clustering in regular client or in-house work, keep a carefully chosen sample of queries and query pairs that people have reviewed.

Include straightforward examples, but focus on distinctions the method might struggle with:

  • paraphrases that express the same thing with different wording;
  • queries about the same topic but with different intent;
  • product variants where one modifier changes the meaning;
  • short, ambiguous queries;
  • brand searches versus generic searches;
  • geographic distinctions;
  • terms that connect two subjects.

Judge these examples against the actual task. For URL mapping, ask: “Could one page reasonably satisfy both queries?” For topic discovery, “Do these belong to the same subject area?” may be the more useful question.

Do not carry a label set over to a different task without checking whether those judgements still mean the right thing.

Review cluster centres and boundaries separately

If you only inspect the most obvious members of a cluster, you can easily come away with too much confidence in it.

Review at least three types of query:

  1. Core members: strong, typical examples that show what the cluster represents.
  2. Boundary members: queries with weak membership, close neighbours in other clusters or assignments that change between runs.
  3. Outliers: queries treated as noise, isolated nodes in a graph or items unusually far from the cluster core.

The boundary cases are often where the clustering method is most likely to change a real SEO decision.

Evaluate cluster labels separately from membership

A useful cluster can still have a misleading name.

Automated naming methods may choose a frequent term, the keyword nearest the cluster’s mathematical centre, or a generated summary. Any of these can leave out an important qualifier or suggest that the group covers more than it does.

Check:

  • whether the label describes most core members;
  • whether it distinguishes the group from nearby clusters;
  • whether it retains important entities, such as brands or products, and meaningful modifiers;
  • whether a practitioner can understand it without opening every row.

A cluster name helps people interpret the result. It is not evidence that the underlying grouping is correct.

Check coverage and fragmentation

Two opposite problems are worth watching for.

Over-merging produces large clusters that combine several distinct intents or page types you would need to handle separately.

Over-fragmentation produces multiple small groups that one useful page or category could reasonably cover.

Useful checks include the spread of cluster sizes, the proportion of noise or single-keyword clusters, each cluster’s nearest neighbours and the number of groups with little separation between them.

Use these measures to flag groups for review, not to trigger automatic fixes.

Add search evidence when the task is page-oriented

If you are making decisions about pages, bring in evidence from search rather than relying only on keyword meaning.

Search engine results page (SERP) overlap shows whether queries currently return many of the same URLs. Search Console shows query and page performance for your own property. An inventory of existing pages can reveal URLs that compete with or complement one another.

These signals answer different questions from semantic similarity. Where possible, record them separately instead of combining them into a single score that hides how the judgement was made.

A practical quality scorecard

A useful review records several dimensions rather than trying to express everything in one number:

  • Cohesion: are the core members meaningfully related?
  • Separation: can you distinguish the group from its neighbours?
  • Stability: does the core survive reasonable parameter changes?
  • Coverage: are important related queries missing?
  • Intent consistency: are the tasks users want to complete compatible?
  • Page-format consistency: could the same type of content plausibly serve the group?
  • Explainability: can a practitioner understand why the group exists?
  • Uncertainty: are ambiguous cases visible rather than hidden?

The aim is not to create a perfect combined score. It is to stop one convenient metric from answering several questions it was never designed to address.

Human review should be targeted, not ceremonial

Manual review takes time and introduces subjectivity. Focus it where someone’s judgement could change the decision, rather than adding a review step just to tick a box.

Prioritise:

  • large, high-value clusters;
  • clusters close to important boundaries in your topic or category structure;
  • low-confidence assignments;
  • unstable groups;
  • clusters that would lead you to create, merge or redirect URLs;
  • regulated or commercially sensitive topics.

Automation can handle the obvious cases. Human attention is most useful at the boundaries and where the consequences matter.

ClusterIQ Conclusion

Keyword-cluster quality has several dimensions. Internal metrics describe cohesion and separation. Stability tests show whether the structure survives reasonable changes. Human judgement checks whether the groups actually support your SEO task.

No single metric can replace all three.

The strongest evaluation process starts with a clear objective, keeps uncertainty visible and checks the result where a mistake would change a page, taxonomy or content decision.

Related ClusterIQ analysis

For metrics that compare one clustering result with another and help check for changes between runs, see Adjusted Rand Index and Normalized Mutual Information.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.