Skip to main content
All articles
Clustering
24 September 2026 7 min read

How to choose a similarity threshold for keyword clustering

A similarity threshold can merge useful topics or fragment them. Learn how to calibrate thresholds against your model, dataset and SEO task instead of copying a fixed number.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Three versions of a keyword similarity graph show a low threshold merging communities, a calibrated threshold preserving useful groups, and a high threshold fragmenting the graph.

A similarity threshold takes a continuous score and turns it into an operational rule. Above it, two keywords might be treated as related. Below it, that relationship gets ignored entirely. Sounds simple enough, but that one number can reshape the entire structure of your keyword-clustering result.

Set it too low and weak relationships start pulling clusters together that shouldn't be together. Set it too high and useful relationships vanish, leaving you with fragments and isolated queries scattered everywhere. The awkward part is that there's no threshold that's universally correct for SEO work.

A defensible threshold has to be calibrated against your model, your dataset, and the actual decision the clusters are meant to support.

Why fixed thresholds are so tempting

Thresholds are easy to describe, which is exactly why people reach for them. A rule like "connect keywords when cosine similarity is at least 0.80" is reproducible, fast, and feels satisfyingly objective.

The problem is that number inherits every assumption baked into the representation underneath it. Cosine similarity measures closeness between vectors. It doesn't attach any universal semantic meaning to a particular score.

As we explain in our guide to cosine similarity for keyword clustering, the exact same numerical threshold can behave very differently across embedding models, languages and datasets.

A threshold is model-specific

Embedding models are trained differently and end up creating different vector spaces. A score of 0.80 from one model isn't guaranteed to represent the same relationship as a 0.80 from another.

Even a model that performs well on general semantic textual similarity benchmarks can behave quite differently on:

  • very short product queries;
  • technical terminology;
  • brand and model names;
  • multilingual data;
  • industry acronyms;
  • queries where a single modifier changes the commercial meaning entirely.

That's exactly why the model name and version should be part of a reproducible clustering configuration. If the model changes, threshold calibration needs revisiting rather than silently carrying over.

A threshold is dataset-specific too

The distribution of similarity scores also depends heavily on the corpus you're working with.

Picture two datasets. The first contains 50,000 queries about home improvement spread across hundreds of categories. The second contains 5,000 queries, all about cordless drills.

That second dataset is semantically much narrower. Plenty of query pairs there will naturally score higher on similarity simply because they share such a tightly defined subject. A threshold that separates topics nicely in the broad dataset can end up far too permissive in the narrow one.

So calibration should start by inspecting the score distribution within your actual dataset, not by importing a number from a benchmark or a tutorial you read somewhere.

A threshold is task-specific as well

The same keyword relationships can support very different tasks depending on what you're trying to do.

If the goal is exploratory topic discovery, a fairly permissive threshold can actually help, since it surfaces broader neighbourhoods and interesting bridge concepts.

If the goal is deciding whether two queries should be mapped to the same commercial landing page, the cost of getting it wrong is much higher. A more conservative relationship rule, combined with intent and page evidence, makes more sense there.

It's the same principle behind the distinction between keyword clustering and topic clustering. The right grouping depends entirely on what happens next with it.

Start with a judgement set

Proper calibration needs labelled examples to work from.

You don't need a perfect research-grade dataset for this. You just need enough representative query pairs to test the behaviour that actually matters to you.

Build a sample containing:

  • obvious paraphrases;
  • lexically similar but meaningfully different queries;
  • same-topic, different-intent pairs;
  • brand or product-model changes;
  • geographic variants;
  • ambiguous pairs;
  • clearly unrelated controls.

Get a practitioner to label them using a small set of explicit categories, such as same operational group, related but separate, and unrelated.

These labels won't remove subjectivity entirely, but they make the decision rule visible and give you a repeatable test set for whenever the model changes down the line.

Inspect the distributions, not just the averages

Once your labelled pairs have similarity scores attached, compare the distributions across categories.

An ideal representation would produce a clean split, same-group pairs scoring high, unrelated pairs scoring low, and ambiguous cases sitting in the middle. Real-world data is rarely that tidy.

The overlap between those distributions is actually the useful part. It tells you exactly where a fixed threshold will force a trade-off.

Raise the threshold and you'll likely reduce false joins while increasing false separations. Lower it and you'll improve recall while risking oversized clusters. There's no way to choose sensibly here without first knowing which type of error costs you more for this particular task.

Watch what the threshold does to the graph

In graph-based clustering, your threshold choice directly reshapes the network you end up with.

Keywords become nodes. Similarity relationships become edges. The threshold decides which candidate edges survive and which get dropped.

As the threshold falls:

  • edge count rises;
  • previously separate regions can become connected;
  • bridge terms gain outsized influence;
  • community boundaries can weaken;
  • large clusters may absorb smaller ones.

Raise the threshold and the opposite happens: the network gets sparser, more keywords end up isolated, and communities can fragment apart.

That's why graph diagnostics matter so much during calibration. Don't just look at an average cluster score. Look at cluster-size distributions, isolated-node rates, component sizes and bridge relationships too.

Nearest neighbours can be an alternative to one global cut-off

A single global threshold isn't the only way to build these relationships.

Another option is to keep the strongest k neighbours for each query, possibly combined with a minimum similarity floor. This stops dense semantic regions from generating an excessive number of edges, while still letting sparser regions hold onto their strongest available relationships.

There are trade-offs, of course. A nearest-neighbour rule can force a query to keep relationships even when none of them are particularly strong. Mutual-neighbour rules, where two queries have to select each other, tend to be more conservative.

The important point is architectural: thresholding is a graph-construction choice you're making, not some fixed law of semantic similarity.

Don't tune purely to an internal clustering metric

Metrics like silhouette score can help describe cohesion and separation mathematically, but optimising a threshold purely to maximise an internal score can produce a result that looks great on paper and is genuinely unhelpful for SEO.

A cluster can have excellent separation in vector space while still mixing incompatible page types together. On the flip side, a commercially useful group might contain linguistic variation that makes it look less mathematically compact than it actually is in practice.

So internal metrics need to sit alongside task-based checks, such as:

  • Does the group make sense to a practitioner looking at it?
  • Would the queries plausibly be satisfied by the same page type?
  • Are important product or entity distinctions preserved?
  • Are bridge terms being hidden from view?
  • Does the result stay reasonably stable under small parameter changes?

Test stability, not just one single run

A good threshold shouldn't produce a completely different topical structure after the tiniest tweak.

One way to test this is to run several nearby values and compare:

  • the number of clusters;
  • the proportion of unassigned or isolated queries;
  • large changes in cluster membership;
  • which clusters split or merge;
  • whether important seed queries stay with their expected neighbours.

If a tiny threshold movement causes wholesale restructuring, that's a sign the dataset contains weak or genuinely ambiguous boundaries. The fix is more likely a different representation, an extra signal, or more human review, rather than hunting endlessly for a magic decimal place.

Use SERP and intent evidence where the decision needs it

Embedding similarity describes language relationships. It doesn't tell you what search engines are actually retrieving right now.

If the operational decision is whether queries should share one URL, bring in evidence closer to that actual decision. Search-result overlap can show whether the same pages currently appear for both queries. Search Console can show you query and page performance for your own property. Manual review can assess whether the required page formats are genuinely compatible.

These signals have limits too. Search results shift with time, location, language, device and context, so treat them as useful evidence, not permanent ground truth.

A repeatable threshold-selection process

  1. Fix the representation. Record the embedding model, preprocessing and similarity function.
  2. Create a labelled sample. Include difficult and commercially meaningful cases, not only obvious pairs.
  3. Score the sample and corpus. Inspect the distributions rather than individual anecdotes.
  4. Test a range of thresholds. Measure false joins, false separations and graph structure.
  5. Review boundary cases. Samples near the proposed cut-off reveal more than the easiest matches.
  6. Evaluate downstream usefulness. Apply page-type, intent, entity or SERP checks where appropriate.
  7. Test stability. Make sure small parameter changes don't create unexplained structural swings.
  8. Version the configuration. Save the chosen threshold alongside the model and dataset assumptions.

ClusterIQ Conclusion

A similarity threshold is an operating parameter, not some fixed SEO constant. Its behaviour depends on the model that produced the scores, the distribution of your dataset, and the cost of different errors in whatever downstream task you're supporting.

The most defensible approach stays empirical. Build a judgement set, inspect the score distributions, test several thresholds, examine what happens to the cluster structure, and validate the result against the real decision it's meant to inform.

If a threshold can't be explained in those terms, its apparent precision is mostly just cosmetic.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.