Cluster stability for SEO: how to tell whether a topic survives small changes
A cluster that disappears after a tiny parameter change deserves caution. Learn how perturbation tests, pairwise agreement and consensus can reveal robust and fragile keyword groups.

Farky Rafiq
Founder of ClusterIQ

A single clustering run can look convincing simply because there is nothing to compare it against. Stability testing asks a harder question: does the important structure survive reasonable changes to the pipeline?
If a topic disappears when the threshold moves from 0.74 to 0.75, or half its keywords jump to a different community under another random seed, that does not automatically mean the cluster is wrong. It does mean the boundary is fragile and deserves more caution before you rely on it.
Stability is different from internal quality
A silhouette score describes cohesion and separation within one single result. Stability compares results across multiple runs. A cluster can have a good internal metric and still be unstable under small changes. Just as often, a genuinely useful SEO group can have fuzzy boundaries while still keeping a stable core.
What should you actually perturb?
Useful stability tests include changing:
- nearby similarity thresholds;
- different random seeds;
- small changes to UMAP parameters;
- HDBSCAN parameter changes;
- graph resolution values;
- alternative embedding models;
- small subsamples or bootstrap samples.
The goal is not to create chaos for its own sake. It is to test sensitivity to choices that are genuinely plausible in practice.
Do not compare cluster numbers directly
Cluster IDs are arbitrary. Cluster 12 in one run might correspond to cluster 4 in the next. Compare membership relationships instead: which pairs repeatedly stay together, which groups show strong membership overlap, which queries frequently switch sides, and which clusters keep splitting or merging.
Core stability often matters more than boundary stability
Topics frequently have a stable centre and unstable edges. For example, "trail running shoes", "trail running trainers" and "off-road running shoes" might stay together in every single run, while a broader term like "running footwear" moves between communities. That is genuinely useful information. The core topic can be trustworthy even when the exact membership at the edges is not.
Pairwise co-membership is a useful abstraction
Across multiple runs, calculate how often two keywords land in the same cluster. If a pair stays together in 19 out of 20 reasonable configurations, that relationship is stable. If they only share a cluster in 9 out of 20, it is far more sensitive. This co-membership matrix can itself become the basis for consensus clustering or confidence reporting.
Consensus clustering can summarise repeated runs
Rather than picking one parameter setting as "the truth", consensus methods combine evidence from several clusterings at once. For SEO, this is useful in exploratory work where the aim is to identify robust topic cores. The risk is added complexity: a consensus result still needs explaining, and it can hide the fact that several materially different structures existed underneath it. Keep the underlying run comparisons accessible rather than just the final summary.
Weight stability by consequence, not by novelty
Not every moving keyword deserves manual review. Prioritise instability when it affects high-volume queries, important commercial terms, cluster labels, URL boundaries, taxonomy changes, or merges and redirects. A low-volume boundary term drifting between two reasonable groups is probably harmless.
Model changes deserve a full comparison, not a quick glance
Switching embedding models is a bigger deal than changing one threshold. Run both models against the same judgement set and compare stable core clusters, new relationships, lost relationships, entity handling, outlier rates and downstream URL decisions. The versioning practices in reproducible keyword clustering make this kind of comparison possible in the first place.
Visual similarity is not the same as stability
Two UMAP plots can look similar while cluster memberships have actually changed a lot underneath. Equally, a visually different projection can preserve many of the same pairwise relationships. Use quantitative membership comparison as your primary evidence, not screenshots.
Build a practical stability report
For each cluster, store: stable core size, membership overlap across runs, the most mobile queries, common split or merge partners, the parameter range tested, and a review status. This gives practitioners a much clearer picture than one generic "confidence score" ever could.
Practitioner principle: stable clusters are not automatically correct, but unstable clusters should not be allowed to make irreversible site decisions without review.
Worked example: stable core, unstable boundary
Imagine a cluster of 120 queries. Across ten reasonable runs, 85 queries always stay together, 20 move occasionally, and 15 regularly switch to a neighbouring cluster. Treating the whole group as simply "stable" or "unstable" loses useful information. The 85-query core can support the cluster identity with confidence. The mobile terms define the boundary that deserves review. If a page decision depends on those mobile terms, your confidence should be lower than if it only concerns the stable core.
Stability can inform how you design the interface
A practitioner-facing system can display stable core members normally while marking mobile terms as boundary cases. That turns an abstract statistical property into something genuinely actionable, without forcing users to understand every clustering metric behind it.
ClusterIQ Conclusion
Cluster stability is one of the strongest checks against overconfidence available to you. By perturbing reasonable parameters, seeds, samples and models, you can tell robust topic cores apart from fragile boundaries. That makes the output far more useful for page mapping and taxonomy, because uncertainty gets attached to exactly the places where it actually exists. The goal is not to make every keyword immovable. It is to know which parts of the structure you can rely on and which parts still need judgement.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Silhouette score for keyword clustering: useful diagnostic, poor SEO verdict

Reproducible keyword clustering: how to make every run explainable
