Skip to main content
All articles
Clustering
1 August 2026 4 min read

Density-based vs centroid-based keyword clustering: two different ideas of what a group is

K-means and HDBSCAN do more than use different algorithms: they define a cluster differently. Learn how centroid and density assumptions change outliers, shape and SEO interpretation.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram comparing K-means and HDBSCAN: every point is assigned to a centroid on the left, while dense irregular regions and unassigned outliers appear on the right.

When you are staring at a spreadsheet of three thousand keywords exported from Ahrefs or Search Console, the goal is usually simple: group these into sensible buckets so you can build a content plan or fix a messy URL structure. In the world of data science, two popular methods, K-means and HDBSCAN, are often suggested as the way to do this. However, they are not just different paths to the same result. They actually have very different ideas about what a group looks like.

K-means tries to organise everything around central points. HDBSCAN, on the other hand, looks for areas where data is crowded and stable, and it is perfectly happy to leave some keywords out if they do not fit. Understanding this difference is far more helpful for your SEO strategy than trying to figure out which one is technically better.

Centroid-based clustering asks which centre is closest

K-means is the classic approach. You tell the tool how many groups you want, and it finds central points (centroids) to anchor those groups. Every single keyword in your list must be assigned to a group, even if it is a bit of an oddball. This is great when you need a clean, complete report where every row has a category, or when you already have a rough idea of how many topics you are dealing with.

Density-based clustering asks where stable dense regions exist

HDBSCAN takes a different view. It looks for clusters that are naturally dense and stay together across different scales. If a keyword is a bit weird or does not strongly belong anywhere, HDBSCAN labels it as noise. This is incredibly useful for keyword research because search data is messy. Topic sizes vary wildly, and sometimes a keyword simply does not belong in your main content pillars. Seeing these as unassigned can be more useful than forcing them into a group where they do not fit.

Cluster shape differs

K-means tends to create tidy, circular groups around its centres. But semantic topics in the real world are rarely that neat. A broad topic might have two very busy sub-sections connected by a thin bridge of related terms. Density-based methods are much better at picking up these irregular, organic shapes that reflect how people actually search.

Every-point assignment is both a strength and a weakness

If you are presenting a final report to a client, having 100 percent coverage looks professional. But during the discovery phase, forcing every keyword into a box can hide the truth. If a query is highly specific or nonsensical, K-means will still give it a home. HDBSCAN is honest enough to say it does not know where it goes. At ClusterIQ, we believe that preserving this distinction helps practitioners make better decisions. Sometimes, unassigned is the most accurate label you can have.

Parameter questions differ

With K-means, your big decision is choosing the number of clusters. With HDBSCAN, you are tweaking settings like minimum cluster size. Neither method removes the need for human expertise; they just change where you apply your judgement during the process.

Worked example: a mixed ecommerce corpus

Imagine you have 5,000 queries for a bathroom retailer. Most keywords fall into big buckets like showers or taps. But you also have a few dozen queries for very specific spare parts. K-means might take those spare parts and shove them into the showers category because that is the closest centre. HDBSCAN might keep the big categories clean and leave the spare parts as noise because there are not enough of them to form a dense group. Neither is wrong, but if those spare parts are high-margin items, you might need a specific manual rule or a separate analysis to handle them, rather than just switching algorithms.

Internal metrics can favour one assumption

The way we measure success matters. Some metrics reward how compact a group is, which naturally makes K-means look like the winner. If you are using density methods, you need to look at different signals, such as how stable a group remains when you change settings. This is why it is vital that internal clustering metrics are chosen to match the specific method you are using.

Graph clustering offers a third definition

There is also graph community detection, which looks at how keywords are connected in a network. This is brilliant when the relationships between terms are the most important factor. You can read more about how these all stack up in our guide on K-means vs HDBSCAN vs graph clustering.

Use several methods as analytical lenses

You do not have to pick just one. A smart workflow involves running a quick K-means baseline, then a density model, and comparing the two. Where they disagree is often where the most interesting SEO insights are hiding. Disagreement between models is a signal, not a failure.

Choose from the downstream cost of mistakes

Think about what happens if a keyword is grouped incorrectly. If a wrong assignment leads to a thousand-pound mistake in a content brief, use a density approach that allows for noise. If you absolutely must categorise every product for a site migration, a centroid approach might be the way to go.

Do not let the algorithm define the product

As an SEO, you should be looking at concepts like confidence and topic relationships, not worrying about the underlying math. ClusterIQ uses these complex methods behind the scenes but presents the results in a way that makes sense for your daily tasks.

Test disagreement explicitly

A great way to improve your data is to look specifically at keywords where K-means and HDBSCAN disagree. Reviewing these outliers can help you decide if you need to adjust your topic count or if you have discovered a new niche you hadn't considered. These edge cases are often the best benchmarks for future work.

Practitioner principle: clustering algorithms do not merely find groups. They define what counts as a group under their assumptions.

ClusterIQ Conclusion

Density and centroid methods represent two different ways of looking at your data. K-means gives you a complete, organised map, while HDBSCAN finds the most reliable hubs of information and filters out the rest. For effective SEO, the best approach is to understand these assumptions and validate your keyword groups against actual business goals and page-level decisions.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.