Agglomerative clustering for SEO: when linkage methods change the topic tree
Agglomerative clustering builds groups from the bottom up, but single, complete, average and Ward linkage can produce very different SEO structures. Learn what changes and why.

Farky Rafiq
Founder of ClusterIQ

Agglomerative clustering is particularly useful when you need to know more than just which group a keyword belongs to. It helps you understand how smaller, specific groups roll up into broader themes. Instead of a flat list of categories, you get a view of how topics relate to one another.
The process starts with every individual keyword acting as its own tiny cluster. The algorithm then repeatedly merges the two "closest" groups together. This creates a hierarchy that maps naturally to how we think about websites: a specific product leads to a subcategory, which sits within a main category. Or, in content terms, a specific long-tail query belongs to a subtopic, which is part of a wider article family.
The crucial detail for any SEO or marketer to grasp is that the definition of "closest" isn't fixed. It depends on something called the linkage rule. If you change the linkage, the entire hierarchy of your keyword research can shift significantly.
What agglomerative clustering does
Think of this as a "bottom-up" approach. Every keyword starts in its own bubble. At each step, the algorithm looks for the two bubbles that are the most similar based on your chosen linkage criteria and joins them. This continues until you either reach a specific number of clusters or the groups become too different to merge further.
In technical tools like Scikit-learn, the AgglomerativeClustering function usually offers four main linkage options: Ward, complete, average, and single. These are essentially different ways of measuring the distance between groups of data.
Single linkage follows the closest pair
Single linkage defines the distance between two clusters based on the two individual items that are nearest to each other across those groups. It is like judging the distance between two cities by measuring the gap between their two closest suburbs.
This method is great at finding elongated, chain-like structures. However, it often suffers from a "chaining" effect. This happens when a series of weak, local relationships joins two concepts that really shouldn't be together. In SEO, this occurs when broad "bridge terms" connect two distinct topic families that have no business being in the same bucket. This is very similar to the structural issues we explore in ClusterIQ's graph bridge analysis.
Complete linkage looks at the furthest pair
Complete linkage does the opposite: it measures the distance between the two most distant members of two groups. It will only merge groups if every member of the new, combined group is relatively close to every other member.
This tends to produce very compact, tight clusters. For an SEO, this is a good safeguard against ending up with messy, "catch-all" groups. The downside is that it can be too strict, splitting up a topic that actually belongs together just because there is some natural variation in how people search for it.
Average linkage takes the middle ground
Average linkage looks at the average distance between all members of one group and all members of another. It is often the "Goldilocks" choice, creating a more balanced hierarchy than the extremes of single or complete linkage.
It is particularly useful when your keyword groups aren't perfectly circular but aren't long, messy chains either. The main thing to remember is that the result is still heavily influenced by how you represented your keywords as numbers (the distance metric) in the first place.
Ward linkage minimises variance growth
Ward linkage is a bit more mathematical. It merges groups in a way that keeps the internal "noise" or variance within the clusters as low as possible. In many tools, this is tied specifically to Euclidean distance.
While this is excellent for numerical data, it isn't always the perfect fit for the "sentence embeddings" we use in SEO, where comparing the angles between vectors (cosine similarity) is often more meaningful. This is why your distance choice and your linkage choice need to be considered as a pair, not in isolation.
Worked example: one keyword family, four hierarchies
Imagine you have a list of a few hundred keywords from Ahrefs or Search Console, including terms like:
- keyword clustering;
- keyword clustering software;
- keyword clustering Python;
- topic clustering;
- semantic clustering;
- content clustering;
- SEO topic map;
- content architecture.
Depending on the linkage you choose, your output changes. Single linkage might lump everything into one giant chain because each term has a logical neighbour. Complete linkage might force a hard split between the technical "Python" terms and the strategic "architecture" terms very early on. Average linkage might give you two neat families, while Ward might create very tight groups based on the specific geometry of the keyword vectors.
None of these is the "correct" SEO answer. They are just different ways of looking at the same data. Your job is to pick the one that makes the most sense for your content plan.
Distance threshold can be more useful than choosing a fixed cluster count
When you are exploring a new niche, telling an algorithm to find exactly 40 clusters is often a guess. A better approach is using a distance threshold. This tells the algorithm to keep merging groups until the difference between them hits a certain limit, then stop.
This allows the data to dictate the number of clusters naturally. ClusterIQ can compare several of these thresholds, showing you which topic branches are stable and "real," rather than forcing you to accept one arbitrary cut of the data.
Dendrograms are useful and easy to overread
A dendrogram is a tree-like diagram that shows the sequence of merges. It is a fantastic way to see which keywords joined together early (meaning they are very similar) and which ones stayed apart until the very end.
However, be careful: the vertical distance in these charts represents mathematical similarity based on your linkage choice. It is not a measure of search volume, keyword difficulty, or business importance.
Hierarchical clustering is particularly useful for taxonomy work
A flat list might tell you that 2,000 keywords from Semrush fall into 60 groups. That is helpful, but a hierarchy shows you that those 60 groups actually sit under 12 broader themes. This is pure gold for building site navigation or a content hub.
When you combine this with ClusterIQ's taxonomy workflow, you can start to see a site structure emerge. Just remember that a mathematical hierarchy still needs a human touch to account for business goals and page types before it becomes a final sitemap.
Use linkage disagreement as evidence
If a group of keywords stays together regardless of whether you use average, complete, or Ward linkage, you can be very confident that it is a distinct, coherent topic. If a group only exists under single linkage because of one weird bridge term, you should probably treat it with caution.
ClusterIQ surfaces these differences, treating cluster stability as a feature rather than hiding the complexity behind a single "black box" result.
Do not compare linkage methods on labels alone
AI-generated labels can be deceptive. Two different clusters might be given the same label by an LLM, making them look identical when the actual keyword lists inside them are quite different. Always look at the actual queries, the entities involved, and the intent of the pages before making a decision. The label is just a shorthand summary.
Where agglomerative clustering fits in ClusterIQ
We use this method primarily as a diagnostic and structural layer. It reveals the parent-child relationships between topics and helps identify where boundaries are fuzzy. While other methods like HDBSCAN are great for finding the "core" of a cluster, agglomerative clustering shows us how the whole map fits together.
Practitioner principle: hierarchical clustering is valuable because it shows how groups combine. The linkage rule defines what “combine” means, so that choice belongs in the evidence.
ClusterIQ Conclusion
Agglomerative clustering gives SEO teams a perspective that flat clustering simply cannot: a multi-level view of topic structure. By comparing different linkages and looking for stable branches, ClusterIQ helps you explore taxonomies and topic maps without pretending there is only one "right" way to organise your site.
A practical linkage-selection experiment for ClusterIQ
Instead of guessing which linkage to use, ClusterIQ allows for controlled comparisons. We can run different linkages on the same set of keyword embeddings and see how they perform against a human benchmark or a set of existing page mappings. The goal isn't to find the "prettiest" tree, but the one that reflects the actual distinctions you need to make in your SEO strategy.
This experiment often reveals local failures. You might find that single linkage works well overall but accidentally merges two high-value, distinct topics. Or complete linkage might be too aggressive, breaking up a perfectly good topic into tiny, unusable fragments. By storing these trade-offs, we can tailor the clustering logic to the specific dataset, whether it is a product-heavy e-commerce site or a broad editorial blog.
Use hierarchy levels as separate product views
The beauty of a hierarchical tree is that it can serve multiple purposes. A high-level view is perfect for reporting to stakeholders about broad "topic share." A mid-level view helps with editorial planning for content hubs. A deep, granular view is ideal for fine-tuning internal linking or identifying specific gaps in your content.
The key is maintaining the connection between these levels. You should be able to zoom in from a category to a subtopic and see exactly which keywords and entities are driving that relationship. This is far more powerful than just exporting a single, flat list of cluster IDs.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Hierarchical clustering for SEO: when a tree is more useful than a flat list of clusters

