Skip to main content
All articles
Clustering
24 September 2026 6 min read

Cosine similarity for keyword clustering: what the score really means

Cosine similarity is useful for comparing keyword embeddings, but the score is not a probability of shared intent. Here is how to interpret and calibrate it.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Two closely aligned vector arrows lead to separate query groups, illustrating that cosine similarity indicates semantic closeness but does not prove shared search intent or the same page.

You have 1,500 keywords to organise into useful groups. Some use different words for much the same thing; others share a topic but need very different pages. Cosine similarity can help you find those relationships, but its scores need careful interpretation.

To use it with text, an embedding model first turns each query into a vector, a list of numbers representing aspects of its meaning. Cosine similarity compares the angle between two vectors rather than their raw size. Vectors pointing in similar directions receive a higher score, which is useful when that direction carries information about meaning.

For SEO, the key distinction is this: a cosine score is evidence of similarity under a particular representation. It is not the probability that two keywords share intent, and it does not automatically mean they belong on the same page.

What cosine similarity actually measures

For two vectors, x and y, cosine similarity is their dot product divided by the product of their lengths. The dot product is calculated by multiplying corresponding values and adding the results. Scikit-learn describes cosine similarity as the L2-normalised dot product. If both vectors have already been normalised to unit length, their dot product is equivalent to cosine similarity.

The geometry asks a straightforward question: how closely do these two vectors point in the same direction? That helps with text because modern embedding models aim to place expressions with related meanings near one another in a space with many dimensions.

Sentence Transformers uses cosine similarity as its default function for comparing the meaning of texts. It also supports dot product, Euclidean distance and Manhattan distance. The original Sentence-BERT work was specifically designed to produce sentence embeddings that could be compared efficiently using cosine similarity.

Why keyword clustering uses cosine similarity

Comparing the words in queries works best when those queries share vocabulary. It becomes less useful when people express similar ideas using different words.

Consider:

  • “cheap running trainers”
  • “affordable running shoes”

A word-overlap method finds only a limited match. A suitable semantic embedding model may place these expressions close together because it represents their broader meaning. Cosine similarity then gives us a numerical measure of how closely their vectors align.

For a list of a few thousand queries, this can make it easier to identify promising relationships without checking every pair by hand. The queries can be converted into embeddings and compared, with the resulting candidate relationships feeding a clustering algorithm, a nearest-neighbour search or a graph.

The benefit is not that cosine similarity “understands intent”. It is that it can reveal relationships that exact matching and simple stemming, which reduces words to their stems, would miss.

A high score does not mean “same page”

This is the most important practical limitation, especially if you are using clusters to plan pages.

Suppose these queries receive a high semantic similarity score:

  • “running shoe reviews”
  • “buy running shoes online”

They share a broad subject, but they suggest different tasks. The first points towards evaluation or comparison. The second is much closer to a transactional shopping task.

An embedding model can capture shared subject matter very effectively while giving less weight to distinctions that matter to an SEO team. Search intent, page type, geography, audience and commercial purpose are separate signals to consider.

This is why ClusterIQ treats similarity as an input to a decision, not the decision itself. The same principle applies to keyword clustering versus topic clustering: a group that fits together mathematically is not necessarily a group that belongs on one page.

There is no universal “good” cosine score

A cut-off, or threshold, such as 0.75 or 0.80 looks precise. But the number has no universal SEO meaning.

Scores depend on several factors, including:

  • the embedding model;
  • how the text is cleaned or otherwise prepared before comparison;
  • the language and subject area;
  • query length;
  • whether the model has been tuned for similarity;
  • the similarity function and normalisation used;
  • the distribution of the dataset being analysed.

You cannot safely assume that 0.80 from one model means the same thing as 0.80 from another. Even with the same model, a useful threshold for a tightly defined product catalogue may differ from one for a broad set of editorial queries.

So the useful question is not “what cosine threshold should SEO use?” It is “what threshold produces acceptable relationships for this dataset, model and task?”

Think in distributions, not isolated numbers

A score is easier to interpret when you know what other scores in the dataset look like. Rather than judging one pair on its own, look at how scores are spread across many pairs.

If nearly every query pair scores between 0.70 and 0.85, a threshold of 0.75 may create a very densely connected network. If most pairs score below 0.40 and only a small proportion rise above 0.70, that same threshold behaves very differently.

To calibrate your threshold, examine:

  1. the spread of scores for random query pairs;
  2. the spread of scores for pairs you have manually confirmed are related;
  3. the spread of scores for queries deliberately chosen to be different;
  4. the number of relationships created at each possible threshold;
  5. how cluster sizes and noise change as you move the threshold.

This makes threshold selection a choice based on evidence from your data, rather than a number borrowed from another workflow.

Cosine similarity in a graph workflow

A graph-based workflow makes the role of cosine similarity easier to see. Each keyword becomes a node, or point in a network. A similarity rule determines which pairs get a connection, called an edge. A community-detection method then looks for groups within that network.

Here, the cosine threshold controls something tangible: graph density, or how many connections the network contains relative to the number possible.

Set the threshold too low and unrelated or weakly related queries may become connected. The network can merge into oversized communities. Set it too high and useful connections disappear, leaving tiny, fragmented clusters and isolated keywords.

That is why choosing a threshold cannot be separated from reviewing the results it produces. The “right” threshold depends partly on whether the resulting structure is useful for your task.

For more background on this network view, see our guide to keyword clustering and graph theory.

Cosine similarity can also hide useful ambiguity

Many keyword-clustering systems turn each relationship into a yes-or-no decision: connected or not connected. Real keyword lists are rarely that tidy.

A query can be close in meaning to several groups. “Shower bath screen”, for example, may relate strongly to bath screens, shower screens and shower-bath product categories. A single similarity score cannot tell you which business or page decision is best.

A good workflow can make that uncertainty visible rather than hiding it. Useful approaches include:

  • showing multiple strongly related neighbours;
  • identifying bridge nodes, or keywords that connect different communities;
  • using soft membership, which allows a query to belong to more than one group, or confidence values;
  • labelling outliers and noise;
  • setting aside uncertain assignments for manual review.

Uncertainty is useful information. Forcing an ambiguous query into a neat group can make the report look cleaner while making the underlying decision worse.

A practical calibration workflow

When you start work on a new keyword dataset, this process gives you a sound basis for your choices:

  1. Choose the representation. Decide which embedding model to use and how you will prepare the text before it reaches the model.
  2. Create a small judgement set. Manually label a few hundred query pairs as clearly related, borderline or clearly different.
  3. Inspect score distributions. Compare the spread of scores in those labelled groups. Look for overlap rather than choosing a threshold in advance.
  4. Test several thresholds. Check their effect on graph density, cluster sizes, outliers and obvious errors.
  5. Review boundary cases. Pairs just above and below your proposed threshold tell you more about its usefulness than the easiest examples.
  6. Validate against the task. If you are mapping clusters to URLs, check that the queries suit the same page format and intent, not just that their meanings are related.
  7. Record the configuration. Save the model, threshold, text-preparation rules and date so you can reproduce the result.

This takes longer than copying someone else's threshold, but it is much easier to explain and defend to a colleague or client.

When cosine similarity is most useful

Cosine similarity is particularly useful when meaning is what you need to compare: finding different ways to express the same idea, discovering related language, proposing connections in a graph or comparing large numbers of query vectors.

It is less useful when the deciding factor is not captured in the representation. If your business decision depends on stock status, geography, regulated terminology, page type or a commercial rule, you need to bring those signals into the workflow separately.

The same distinction applies to search behaviour. Similarity between query embeddings and similarity between observed search results measure different things. Neither should be treated as a straightforward substitute for the other.

ClusterIQ Conclusion

Cosine similarity gives keyword clustering a consistent mathematical way to compare query vectors. Paired with semantic embeddings, it can help you find related expressions even when they share few words.

The mistake is treating its score as a universal measure of SEO intent. What a score means depends on the model, dataset and task. Use it as evidence, calibrate it against your own data and check the effect on the clusters and decisions that follow.

The goal is not the highest possible similarity score. It is a set of relationships that helps you make better decisions about your keywords and pages.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.