HDBSCAN for keyword clustering: density, noise and uncertain queries
HDBSCAN can discover variable-sized keyword groups and leave uncertain queries as noise. Learn how its parameters work and where the method fits SEO clustering.

Farky Rafiq
Founder of ClusterIQ

HDBSCAN appeals to a lot of people doing keyword clustering because you don't have to decide the number of clusters up front, and it's willing to leave some queries unassigned rather than forcing them somewhere they don't belong. Those two traits fit a reality that SEO datasets throw at you constantly: topical groups come in wildly different sizes, and some queries genuinely sit outside the main structure altogether.
That doesn't make HDBSCAN the automatic right answer, though. Density-based clustering carries its own assumptions, parameters and failure modes. The quality of what comes out still depends heavily on how you represented the keywords and which distance relationships you preserved along the way.
What HDBSCAN actually does
HDBSCAN takes density-based clustering and turns it into a hierarchical process. The reference implementation describes the method as building a hierarchy first, then extracting a flat clustering based on how stable each cluster turns out to be.
Instead of asking you for a fixed number of groups, the algorithm looks for regions of the data that stay meaningfully dense across different distance scales. Points that don't belong strongly enough to any of those regions get labelled as noise.
For keyword clustering, that's genuinely useful, because your dataset doesn't get forced into one complete partition. A query that doesn't fit any stable semantic neighbourhood can just sit outside the main clusters, ready for you to look at later.
Why noise is a feature, not a failure
A lot of SEO clustering workflows quietly assume every single keyword has to belong somewhere. That's operationally convenient, but it can hide a great deal of genuine uncertainty.
A large keyword export typically contains:
- irrelevant queries;
- very specific long-tail expressions;
- brand terms that behave differently from generic language;
- ambiguous queries;
- terms that bridge two topics at once;
- isolated product or technical phrases.
Force each one of those into its nearest cluster and your dataset looks complete on paper, while quietly hiding a load of weak assignments underneath.
HDBSCAN's ability to mark points as noise gives you an explicit review surface instead. An unassigned query isn't necessarily unimportant, it just means that point didn't meet the density conditions needed for stable cluster membership.
The role of min_cluster_size
The main HDBSCAN documentation describes min_cluster_size as the smallest grouping you're willing to consider a cluster. That's fairly intuitive on its own, but its effect isn't isolated from everything else.
Set min_cluster_size very small and the algorithm will hang on to lots of small groups. Set it larger and smaller structures can disappear entirely or get relabelled as noise.
There's no universally correct value for an SEO dataset. A product catalogue might contain perfectly legitimate three-query micro-topics around specialist parts, while a broad editorial analysis might treat a group that small as too unstable to act on.
So this parameter is as much an operational choice as a mathematical one: how small can a group get before it stops being useful for the task at hand?
min_samples changes how conservative the clustering is
HDBSCAN also exposes min_samples. In the reference implementation, if you don't specify it, it defaults to whatever value you gave min_cluster_size. The documentation notes that larger values make the clustering more conservative overall, meaning more points get labelled noise and clusters get restricted to denser regions.
This interaction catches people out. You might change min_cluster_size expecting to alter only the smallest permissible cluster, while unintentionally also changing how conservative the whole result is, because min_samples was still coupled to it.
For reproducible SEO analysis, record both values explicitly rather than leaving one to default silently.
Representation still matters more than the algorithm's name
HDBSCAN doesn't cluster raw meaning. It clusters points under whatever representation and distance metric you feed it.
Feed it TF-IDF and density exists in a lexical feature space. Feed it sentence embeddings and density exists in the model's semantic vector space instead. If you've applied dimensionality reduction first, the algorithm only ever sees that reduced representation, not the original one.
That means the exact same keyword list can produce noticeably different HDBSCAN clusters depending on:
- the embedding model;
- normalisation;
- distance metric;
- dimensionality reduction;
min_cluster_size;min_samples;- the cluster selection method.
A sensible workflow treats the whole configuration as part of the result, not just an implementation detail to skip past.
Be careful when reducing dimensions first
Keyword embeddings can run to hundreds of dimensions. It's common to reduce that space before density-based clustering, either for speed or to sharpen local structure.
UMAP gets used for this a lot. Worth remembering that it's a dimensionality-reduction technique, not a clustering algorithm in its own right. Its original paper presents it as a scalable manifold-learning method useful for visualisation and general-purpose dimension reduction.
The practical caution here is simple: run HDBSCAN on a UMAP projection and your clustering is based on relationships that exist after that transformation. The projection's own parameters have effectively become part of your clustering pipeline, whether you meant them to or not.
A two-dimensional visualisation is a particularly risky clustering input, mostly because it's so easy to plot. A representation built for human display isn't automatically the best representation for making cluster decisions.
Soft membership can expose uncertainty
The HDBSCAN implementation can give you membership strengths for clustered points, which is useful because it means cluster assignment doesn't have to be treated as equally certain for every keyword.
In practice, you can route lower-confidence assignments for human review while letting strong core members pass through automatically. That's a much more honest workflow than pretending every cluster boundary is perfectly crisp.
For SEO purposes, the uncertain queries often contain exactly the information you care about most. They can reveal:
- overlapping search tasks;
- bridge concepts between categories;
- ambiguous language;
- missing taxonomy structure;
- cases where a separate page may genuinely be needed.
The goal isn't to eliminate these cases entirely. It's to make them visible so someone can look at them.
When HDBSCAN fits keyword data well
HDBSCAN tends to be a good fit when:
- you don't know the appropriate number of clusters in advance;
- cluster sizes are expected to vary a lot;
- you want the system to leave weak points unassigned rather than forced;
- you care more about stable dense regions than forcing complete coverage;
- the representation produces genuinely meaningful local neighbourhoods.
It works particularly well as an exploratory layer over semantic embeddings, especially when the aim is to surface candidate groups for a practitioner to review rather than acting on them automatically.
When HDBSCAN may be a poor fit
It's less attractive when every item needs a deterministic business category, when legitimate groups are extremely sparse, or when density varies so wildly across the dataset that one parameterisation can't represent the whole thing well.
It can also be the wrong tool when the relationships are better modelled as a graph. A keyword can have several meaningful connections across topics at once. Community detection can preserve that network structure more naturally than clustering points purely by local density.
That's one reason ClusterIQ treats clustering methods as interchangeable components rather than one single definition of truth. Our graph theory guide covers the alternative network view in more depth.
HDBSCAN doesn't solve search intent
A dense semantic group can still contain several different user tasks underneath it.
"Best CRM software", "CRM pricing" and "how CRM software works" can all sit in the same broad semantic region while genuinely needing different page formats. HDBSCAN can tell you the language is related. It can't tell you the right information architecture on its own, without more evidence.
That's consistent with the broader principle covered in keyword clustering versus topic clustering: a coherent group is evidence for a decision, not the decision itself.
A practical HDBSCAN evaluation workflow
- Fix the representation. Record the embedding or lexical features and the distance assumptions.
- Start with a plausible minimum cluster size. Base it on the smallest grouping that could be operationally useful.
- Set
min_samplesexplicitly. Don't let it change accidentally while you're testing cluster size. - Inspect the noise rate. Check whether unassigned queries are genuinely ambiguous, or whether useful groups are being lost.
- Inspect cluster-size distribution. Watch for one giant cluster or a mass of trivial micro-clusters.
- Review membership confidence. Sample strong core points and weak boundary points separately.
- Test nearby parameter values. Stable topics shouldn't vanish after tiny changes for no explainable reason.
- Validate against the SEO task. Add intent, page type, SERP or business rules before you act on the output for site architecture.
ClusterIQ Conclusion
HDBSCAN offers a genuinely useful compromise for keyword clustering. It can discover dense groups of varying sizes without needing a fixed number of clusters up front, and it can keep uncertain observations outside those groups instead of forcing them into a bucket where they don't quite belong.
Its strength isn't that it removes judgement from the process. Its strength is that it represents uncertainty more explicitly than most alternatives.
Treat the noise points, membership strengths and parameter sensitivity as diagnostic information. If the end goal is an SEO decision, validate the mathematical structure against intent, page format and site purpose before you act on it.
Related ClusterIQ analysis
For the closest density-based comparison, see DBSCAN vs HDBSCAN for keyword clustering.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Adding new keywords to existing clusters without rebuilding everything

Soft clustering and confidence scores: handling ambiguous keywords honestly
