DBSCAN vs HDBSCAN for keyword clustering: fixed density or adaptive hierarchy?
DBSCAN and HDBSCAN both find dense regions and allow noise, but they make different assumptions about density and parameter choice. Learn what that means for SEO datasets.

Farky Rafiq
Founder of ClusterIQ

DBSCAN and HDBSCAN are frequently mentioned in the same breath because they both rely on density to group data. Unlike K-Means, which forces every point into a cluster, these methods are comfortable leaving outliers behind as noise. This is a massive advantage for SEOs who do not want to force every weird, long-tail query into a category where it does not belong.
The practical difference, however, is significant. DBSCAN expects your data to follow one specific rule for density across the entire set. HDBSCAN is more flexible, looking for clusters that remain stable even as density levels change. Because keyword data is rarely uniform, understanding which one to use is vital for building accurate content plans.
What DBSCAN does
DBSCAN works by finding core points that have a minimum number of neighbours within a specific radius, known as eps. If a point has enough neighbours, a cluster starts to grow. Any keyword that cannot be reached from one of these dense regions is simply labelled as noise.
The two main levers you have are:
eps: the maximum distance between two points for them to be considered neighbours.min_samples: the minimum number of points required to form a dense region.
One of the best things about this approach is that you do not need to tell the algorithm how many clusters to find beforehand. It discovers them based on the density of your keyword list.
The challenge of a single global radius
DBSCAN performs beautifully when your clusters all have a similar density. In the real world of search data, this is rarely the case. If you are looking at a few thousand keywords from Ahrefs or Search Console, you will notice that broad commercial topics are often incredibly dense, with hundreds of slight variations of the same phrase. Meanwhile, technical or niche subtopics might only have ten or fifteen keywords that are semantically related but phrased quite differently.
This creates a dilemma. If you set a tight eps to keep your broad topics distinct, your niche topics might be discarded as noise. If you loosen the radius to capture those smaller groups, your broad topics might merge into one giant, unhelpful blob. This is a common frustration when trying to automate URL mapping for a large site.
HDBSCAN removes the need for one fixed radius
HDBSCAN solves this by looking at how groups persist across different density levels. Instead of forcing you to pick one perfect eps value, it builds a hierarchy of possible clusters and then extracts the ones that are most stable.
This adaptability is why HDBSCAN is attractive for keyword clustering. It can handle a dataset where one cluster is a tightly packed group of "running shoes" queries and another is a sparse, spread-out group of "orthopaedic trail running footwear for flat feet" queries. It recognises both as valid groups because they are internally consistent, even if their densities differ.
Noise exists in both methods
Both algorithms allow for a "noise" category, which is incredibly useful for SEO. Not every query in a Search Console export deserves its own page or even a mention in a content brief. By allowing keywords to remain unassigned, you avoid polluting your clusters with irrelevant data.
The noise set often contains:
- Completely irrelevant or "junk" keywords.
- Unique, one-off entities that do not belong to a broader topic.
- Emerging niche opportunities that haven't reached critical mass yet.
- Queries that sit awkwardly between two distinct topics.
Rather than just deleting these, ClusterIQ treats them as a review set. Sometimes the "noise" is where the most interesting new content ideas are hiding.
Worked example: mixed product demand
Consider a dataset of 40,000 keywords for a large bathroom retailer. You have massive categories like "showers" and "taps" which are semantically dense. You also have highly specific "spare parts" queries for obscure European brands.
A global DBSCAN radius might struggle here. It might correctly group the showers but then fail to see the spare parts as a cluster, or it might group all the spare parts together regardless of brand just to make them fit. HDBSCAN is much better at preserving these variable-density groups, though it still relies on your embeddings being high quality enough to distinguish between the concepts in the first place.
DBSCAN can still be useful
HDBSCAN is not always the superior choice. There are times when the simplicity of DBSCAN is actually a benefit. It can be valuable when:
- Your dataset has a very clear, uniform density.
- You want a simple, transparent baseline to compare other models against.
- You have a specific "neighbourhood" distance in mind that you can calibrate using known examples.
If a simpler model gives you the clusters you need to build a solid content plan, there is no need to overcomplicate things.
Parameter sensitivity differs
DBSCAN is notoriously sensitive to the eps setting. A tiny change can cause a cluster to explode into noise or merge with its neighbour. HDBSCAN shifts the focus to min_cluster_size. This is often more intuitive for a marketer: you are essentially telling the tool, "I don't care about any topic that has fewer than five keywords."
ClusterIQ looks at how stable your clusters are across different settings, rather than just giving you one take-it-or-leave-it result. This helps you see if your clusters are robust or just a fluke of the settings.
Distance metrics still matter
No clustering algorithm can fix bad data. Both methods rely on the underlying geometry of your keyword embeddings. If your embedding model cannot tell the difference between "apple the fruit" and "Apple the tech company," no amount of density analysis will fix that. Your choice of algorithm should always be paired with careful embedding evaluation to ensure the machine sees the same relationships a human SEO would.
DBSCAN does not solve the high-dimensional problem
When you work with high-dimensional data, distances between points start to look very similar, which makes density hard to define. Sometimes we use dimensionality reduction like UMAP or PCA to help. If you do this, remember that the reduction itself changes the "shape" of your data. It should be tested as part of the whole workflow, not just as a pre-processing step.
HDBSCAN offers richer membership evidence
One of the best features of HDBSCAN is that it provides a "probability" or membership strength for each keyword. This tells you how strongly a keyword belongs to its assigned cluster.
In ClusterIQ, this allows for better review queues. Instead of just seeing a list of keywords, you can see which ones are the "core" of the topic and which ones are on the fringes. Our guide on soft clustering and confidence explains how to use this data to make better decisions without getting bogged down in complex statistics.
Use disagreement as a review set
A smart way to work is to run both methods on the same data. If both DBSCAN and HDBSCAN agree that a group of keywords belongs together, you can be very confident in that cluster. If they disagree, those keywords are perfect candidates for manual review. This disagreement often highlights areas where your keyword research is ambiguous or where a new subtopic is starting to form.
Do not choose from cluster count alone
It is tempting to think that more clusters equals more detail, but that is not always true. A model that gives you 200 tiny, fragmented clusters might be less useful than one that gives you 80 solid, actionable topics. Focus on how well the clusters translate into page-level decisions and content briefs.
Practitioner principle: DBSCAN assumes there is one useful scale for density. HDBSCAN searches across multiple scales. The right choice depends on whether your keywords are tightly packed or spread out across different niches.
ClusterIQ Conclusion
DBSCAN and HDBSCAN are both excellent tools for the SEO who wants to avoid forcing data into boxes. While DBSCAN is a great baseline for uniform data, HDBSCAN is usually the better fit for the messy, variable-density world of search queries. By using these tools to identify noise and stable clusters, you can build content strategies that are based on the actual structure of the data, not just guesswork.
A parameter-selection workflow that exposes the trade-off
When using DBSCAN, we can plot the distances to the nearest neighbours to find a sensible eps. For HDBSCAN, we look at how the noise rate and cluster stability change as we adjust the minimum cluster size. The goal isn't to eliminate all noise or find the most clusters possible. The goal is to find a "stable" region where your high-value decisions, like which keywords belong on a specific landing page, stop shifting around.
ClusterIQ shows you these trade-offs directly. You can see coverage, stability, and noise side by side. This makes the technical side of SEO understandable for clients and stakeholders without losing the underlying rigour.
Use noise as a product workflow
The "noise" category should not be a dead end. In a good workflow, these unassigned keywords are routed into specific buckets for review: perhaps they are new niche topics, ambiguous terms, or just junk. Over time, as you add more data, these outliers might form their own clusters. This turns "unassigned" into a temporary state, allowing your topic map to grow and evolve naturally as you gather more search data.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Density-based vs centroid-based keyword clustering: two different ideas of what a group is

Soft clustering and confidence scores: handling ambiguous keywords honestly
