Skip to main content
All articles
Clustering
4 August 2026 4 min read

Silhouette score for keyword clustering: useful diagnostic, poor SEO verdict

Silhouette score can describe separation and cohesion in a clustering result, but it cannot tell you whether the groups make sense for search intent, pages or business decisions.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Side-by-side keyword clusters showing a tidier geometric grouping with mixed intent versus a slightly overlapping grouping that separates pricing and comparison queries more usefully for SEO pages.

The silhouette score is a tempting metric because it boils down a complex clustering project into a single, digestible number. When you are staring at a spreadsheet of three thousand keywords from Ahrefs or Search Console, having one figure to tell you if the groups are "good" feels like a massive time-saver.

However, while it is a brilliant diagnostic tool for experiments, it is a risky way to make final SEO decisions. A high score suggests your data points are tucked neatly into their own groups and kept far away from their neighbours. What it cannot do is tell you if those groups actually represent logical SEO topics, matching search intents, or sensible page structures.

What silhouette score actually measures

For every keyword in your set, the silhouette calculation compares two things: how close that keyword is to others in its own cluster, and how far it is from keywords in the next nearest cluster. Scikit-learn, a popular tool for this kind of data science, calculates a coefficient based on these internal and external distances.

The results sit between -1 and 1. A high positive value usually means strong separation. Scores near zero suggest the boundaries are messy and overlapping, while negative scores might mean a keyword has ended up in the wrong group entirely.

Why it is useful for SEOs

Silhouette is at its best when you are comparing different technical setups using the same data. For example, ClusterIQ might run several tests using HDBSCAN or different values for a K-means baseline. In this scenario, the score helps us spot configurations that create weak geometry or highlight specific groups where keywords are sitting uncomfortably close to a different cluster.

Why it is not an SEO metric

The fundamental issue is that silhouette knows nothing about the real world of search. It has no concept of intent, page types, brand nuances, or your existing URL structure.

If your embedding model decides that "CRM software pricing", "best CRM software", and "how CRM works" are all semantically similar, they will be grouped tightly together. This might result in a fantastic silhouette score, but for a content plan, it is a disaster. Those three queries represent different stages of the funnel and usually require three distinct pages.

The data determines what the metric sees

Because silhouette is calculated from distances, it is only as good as the data you feed it. If the initial representation of your keywords is flawed, the metric will simply validate that flaw.

This is why comparing a TF-IDF approach against a semantic embedding approach using only silhouette scores is often a mistake. They are looking at different types of evidence. Our guide on TF-IDF versus semantic embeddings goes into more detail on why this distinction matters for your keyword maps.

The choice of distance matters

You can calculate silhouette using various distance metrics. When dealing with text embeddings, cosine distance often makes more sense than Euclidean distance. If the way you measure the score differs from the way the clusters were actually built, the resulting number won't accurately reflect the structure of your keyword groups.

Noise and singletons

Modern density-based clustering often produces "noise" (keywords that don't fit anywhere) or tiny clusters of just two or three terms. These can make your average scores fluctuate wildly. Before you trust a headline figure, you need to know if the noise was excluded and whether the score is a global average or broken down by individual cluster.

Per-cluster silhouette is often more useful

A single average score can hide a lot of problems. You might have nine clusters that are perfectly separated, but one commercially vital cluster that is a tangled mess with its neighbours. The overall average might still look great, masking the fact that your most important content brief is based on poor data.

In ClusterIQ, we find that looking at local diagnostics is far more helpful. It allows a practitioner to ignore the "perfect" groups and focus their manual review on the specific clusters that the metric has flagged as weak.

A practical example: the "wrong" winner

Consider two different ways of clustering a thousand keywords. Configuration A gives you a high silhouette score of 0.51 but lumps pricing, comparisons, and category terms into one giant "commercial" bucket. Configuration B scores lower at 0.46 but successfully separates pricing from comparisons because it recognises the SERP intent is different.

If you are trying to decide which URLs to build, Configuration B is the clear winner for SEO, even though the "maths" says Configuration A is more efficient.

Using silhouette in a scorecard

Rather than relying on one number, a professional evaluation should look at a range of factors: how stable the clusters are when parameters change, how well they match human judgement, and whether they respect intent and page types.

This is the core of ClusterIQ's cluster-quality approach. We believe a metric should describe a specific property of the data rather than trying to provide a final "pass or fail" verdict on your SEO strategy.

Don't optimise for the score alone

If you or your team focus purely on chasing a higher silhouette score, you might end up with a model that is great at geometry but terrible at SEO. This often leads to clusters that are too broad, the removal of "messy" but valuable long-tail queries, and a content plan that ignores the complexity of the real SERPs.

The goal is a useful structure for your website, not a perfect score in a notebook.

How we use silhouette in ClusterIQ

We treat this metric as helpful context. It should be visible as a diagnostic for the whole run and for individual clusters, acting as a warning system for weak groups. It belongs alongside plain-English explanations and manual controls, not as an opaque "quality score" that you aren't allowed to question.

Practitioner principle: Silhouette can tell you if a cluster is mathematically distinct in its data space. It cannot tell you if those keywords belong on the same page.

ClusterIQ Conclusion

The silhouette score is a valuable engineering tool for keyword clustering. It is excellent for comparing different technical setups and finding regions of your data that need a closer look.

However, its role should be specific and narrow. Always weigh it against stability, search intent, and human expertise before you use a cluster to change your site's architecture.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.