Hierarchical clustering for SEO: when a tree is more useful than a flat list of clusters
Flat clusters hide relationships between broad topics and subtopics. Hierarchical clustering can expose that structure, but dendrograms still need practical SEO interpretation.

Farky Rafiq
Founder of ClusterIQ

Flat keyword clusters give every query one cluster ID, and every cluster sits next to the others as if they're all equally important. That's handy, but it's not how topics actually work.
Real topics nest inside each other.
Take "running shoes". Inside that you've got trail shoes, stability shoes, racing shoes and beginner shoes. Each of those can split further into narrower groups of products, questions and use cases. Hierarchical clustering is useful here because it captures relationships at several levels at once, rather than forcing everything into one fixed level of detail.
What hierarchical clustering changes
A flat clustering answers one question:
Which group does this query belong to?
A hierarchy adds a second question:
How do the groups relate to each other at broader and narrower levels?
That extra layer makes the result more useful for:
- topic maps;
- site taxonomy;
- content hubs;
- cluster review;
- picking the right level of granularity for the job at hand.
Agglomerative clustering builds upward
Scikit-learn's AgglomerativeClustering works by repeatedly merging pairs of clusters using a linkage rule you choose.
It starts with individual keywords, or small groups of them, and gradually builds larger and larger groups as it merges.
The common linkage choices are:
- Ward: merges groups to keep within-cluster variance as low as possible;
- complete: looks at the maximum distance between members of two groups;
- average: uses the average pairwise distance between groups;
- single: uses the smallest pairwise distance between groups.
Each one produces a different shape of tree, and each has different weaknesses around chaining and outliers.
Why the dendrogram is so useful: it shows you every possible cut
A dendrogram is simply a picture of the merge history.
Rather than deciding upfront that your dataset contains exactly 40 topics, you can look at how groups combine as you relax the distance threshold.
At one cut, you might see:
- trail running shoes;
- stability running shoes;
- racing shoes.
At a broader cut, all three merge into "running shoes".
This matters because SEO work happens at different levels. A navigation system usually needs the broad category, while a content brief needs the narrower topic.
A hierarchy is not automatically a taxonomy
A dendrogram is a statistical structure, not a finished site hierarchy.
A parent-child merge can make perfect mathematical sense while still being awkward for navigation. Two groups might sit close together semantically but belong under completely different commercial sections.
Treat the hierarchy as evidence of topical closeness, then check it against:
- entities;
- page type;
- business taxonomy;
- user journeys;
- inventory or content depth.
Distance choice still matters
Hierarchical clustering doesn't let you dodge the representation question.
If the inputs are semantic embeddings, the tree reflects that embedding space. If the inputs are TF-IDF features, it reflects lexical similarity instead.
Cosine distance can work well for text, but the linkage method you choose needs to support that metric, and the vectors need to be prepared correctly first.
Our comparison of TF-IDF and semantic embeddings covers how each representation preserves different kinds of evidence.
Single linkage can create chains
Single linkage joins two clusters as soon as any single pair of points is close enough. That makes it prone to chaining, where a run of in-between observations bridges topics that are otherwise quite different.
With keyword data, a sequence of bridge terms can end up connecting topics that have little in common overall.
That can be interesting to explore, but it can also produce broad, stretched-out clusters that aren't much use operationally.
Complete and average linkage tend to be easier to interpret
Complete linkage is more cautious because it looks at the furthest pair between two groups. Average linkage sits in the middle, using the average distance instead.
There's no single winner here. Test each structure against a judgement set and against the task you actually need it for.
For topic exploration, a few broad bridges might be fine. For page mapping, tighter groups are usually more useful.
Hierarchies make cluster naming easier
When a narrow cluster sits under a clear parent, naming gets much simpler.
For example:
- parent: keyword clustering algorithms;
- child: density-based clustering;
- grandchild: HDBSCAN parameter selection.
The parent supplies the context, so the child label doesn't need to spell out the whole topic every time.
This pairs well with the approach in our cluster-naming guide.
Hierarchy can expose over-fragmentation too
If ten tiny clusters merge almost immediately at a very small distance, they may just be minor variations rather than genuinely useful distinctions.
On the flip side, if two clusters stay separate until very late in the tree, forcing them together for convenience risks hiding a real boundary.
That makes the hierarchy a handy review tool for spotting whether your flat result is too coarse, or too fragmented.
A practical SEO workflow
- Create a stable representation of your queries.
- Run agglomerative clustering with a small set of plausible linkage methods.
- Look at the dendrogram or the merge distances.
- Compare broad cuts against narrow cuts.
- Label the recurring parent and child concepts.
- Check the proposed hierarchy against entities, page types and site structure.
- Use the hierarchy as a topic map, not as an automatic URL tree.
Practitioner principle: hierarchical clustering earns its keep by showing you that "the right number of clusters" depends entirely on which decision you're making.
ClusterIQ Conclusion
Hierarchical clustering gives SEO teams a single model that shows both broad topics and narrow subtopics.
Agglomerative methods and dendrograms reveal where groups merge and how stable different levels of granularity are. That's genuinely useful for topic maps, taxonomies and content planning.
The hierarchy still needs a human eye on it. Statistical containment is evidence of topical structure, not permission to copy the tree straight into your site's URLs.
Related ClusterIQ analysis
For a deeper look at how the hierarchy itself is built, see agglomerative clustering and linkage methods.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Agglomerative clustering for SEO: when linkage methods change the topic tree

Soft clustering and confidence scores: handling ambiguous keywords honestly
