When semantic similarity helps keyword clustering, and when it does not
Semantic similarity can uncover related queries that lexical matching misses, but it cannot decide whether those queries belong on the same page. Learn how to combine embeddings with search intent, SERP evidence and human review.

Farky Rafiq
Founder of ClusterIQ

If you have a spreadsheet of 1,500 keywords and need to turn it into a content plan, putting similar phrases together is only part of the job. You still need to decide what each group means for the site: one new URL, one page brief, several pages or a particular content format.
Semantic similarity can help you find connections that matching words alone would miss. It can pick up paraphrases, synonyms and related expressions, even when the queries use quite different wording. But two searches being close in meaning does not prove that people want the same thing, or that one page should serve both.
The practical approach is to use semantic similarity to find possible relationships, then check them against search intent, search results, business context and human judgement. A useful keyword cluster is a decision neighbourhood: a group that supports a shared action, not just a group of related phrases.
What semantic similarity adds to keyword clustering
Keyword grouping often starts by looking for shared words, word roots, modifiers or named things such as brands and products. These are known as lexical signals, and they still matter. They can identify near-duplicates, spelling variations, product names, locations and model numbers with a level of precision that a general language model may not provide.
The catch is that people can describe closely related things using different words. Take these illustrative queries:
adjustable standing desk
height adjustable sit stand desk
electric desk that raises and lowers
A method based on shared words, sometimes called token overlap, may put the third query further away from the first two. A semantic model can instead turn each query into a numerical representation called a vector, then measure how close those vectors are. That gives it a way to find relationships based on meaning rather than shared wording. How much this helps still needs testing on your own keyword dataset.
Word embeddings represent individual words as continuous numerical vectors learned from patterns of language use. Sentence-embedding models apply the same broad idea to phrases or complete queries. Sentence-BERT, for example, was designed to produce sentence representations that can be compared efficiently, including for semantic search and clustering. In practice, the model converts each query into numbers that you can compare or group.
One common way to compare them is cosine similarity. This measures how closely two vectors point in the same direction, rather than simply measuring the distance between their endpoints. A higher score generally means their directions are more similar within that model's vector space.
The score can help rank possible matches or flag queries worth reviewing together. It is not a universal test of whether two keywords belong on one page. A cosine score of 0.8 from one model, language or dataset cannot automatically be treated as equivalent to 0.8 from another.
Semantic similarity is not search intent
Search intent is about what someone is trying to do. Semantic similarity is about how a model represents the relationship between pieces of language. The two overlap, but they answer different questions.
For example, these queries may be semantically close:
best standing desk
standing desk price
how to choose a standing desk
They all concern the same product, but the likely tasks differ. One person wants to compare options, another wants commercial pricing information and a third may want buying advice. Depending on your site and the search results, those needs could call for different page types.
Shared wording can also hide distinctions that matter:
standing desk converter
electric standing desk
Both include “standing desk”, but they describe different products. If a similarity method puts too much weight on that shared phrase, it could merge queries that should stay separate.
Short queries are particularly tricky. Someone searching for apple watch might want a product, a retailer, support information or a comparison. The query itself may not provide enough context for a model to work out the task reliably.
At ClusterIQ, we treat semantic similarity as a useful signal, not proof that two queries belong on the same URL. The difficult judgement is usually deciding when related language also points to a shared SEO action.
Why the clustering method still matters
Embeddings do not create clusters on their own. They give you representations of queries, or relationships between pairs of queries. You still need a separate method to turn those relationships into groups.
K-means
K-means assigns queries to a chosen number of groups, each organised around a centre point called a centroid. It can be useful when the data forms reasonably compact groups and you know, or can estimate, how many clusters you need. That requirement becomes less convenient when you do not yet know the structure of your keyword list.
Hierarchical clustering
Hierarchical clustering can show groups within groups. A broad topic might contain separate product, comparison and informational subgroups. The result depends on how distance is measured, how groups are linked together and where you cut the hierarchy to choose your final clusters. A useful topic hierarchy is therefore not automatically a ready-made site architecture.
Density-based clustering
Methods such as HDBSCAN look for dense groups and can leave some queries outside them as outliers. This can help when a keyword list contains clear themes alongside weak connections and unusual searches. The results still depend on assumptions about density and settings such as minimum samples. Changing a setting can split groups, merge them or mark parts of the same list as noise rather than assigning them to a cluster.
Graph-based approaches
A keyword graph treats each query as a node and each relationship as a connection, or edge, with a weight attached. Community detection then looks for groups within that network. This can help you inspect themes, queries that connect themes and outliers. But the result depends on how you create the connections, which threshold you use and how the method identifies communities.
Each algorithm answers a mathematical question about the way your queries are represented. None independently decides whether a group needs one URL, several URLs or a different content format.
A practical hybrid workflow
A reliable workflow gives each signal a clear job. Rather than expecting one method to turn 1,500 keywords into a finished plan, use different checks to build and test your decisions.
1. Clean and normalise the dataset
Remove exact duplicates and decide how you will handle spelling variants, punctuation, locations, brands, product names and model numbers. Keep the original queries available so you can check them later. Tidying too aggressively can remove distinctions that matter when you decide which page should target a query.
2. Use lexical rules for obvious relationships
Group or flag clear near-duplicates and known entities, such as named products or places. Exact wording is especially useful when it carries commercial meaning. A model number, destination, product type or local area may be a reason to keep queries separate, even when the surrounding language is similar.
3. Use semantic similarity to generate candidates
Convert the remaining queries into representations using a suitable sentence-embedding model, then compare them. At this stage, the aim is recall: finding possible relationships you might otherwise miss. That includes paraphrases, links between topics and less obvious connections that deserve a closer look.
Do not treat a similarity threshold as a universal rule. The model, language, query length, preprocessing and specialist terminology all affect the scores. Test your thresholds against examples from your own dataset that reviewers have already labelled.
4. Classify the likely intent
Look at the task behind each query. You might use categories such as informational, comparison, transactional, navigational, local or support intent. The right set of categories depends on the project, rather than on a fixed template.
Modifiers can be important clues. Words and phrases such as best, price, near me, how to and for beginners may point to different needs. They do not always mean you need separate pages, but you should not ignore them just because the main topic matches.
5. Compare the search results
Search engine results page, or SERP, overlap compares the URLs returned for different queries. If two queries repeatedly bring back many of the same relevant pages, that is search-specific evidence that a similar page or result type may serve both. It is often more directly useful for mapping keywords to pages than generic semantic similarity.
Still, SERP overlap is not a definitive answer. Results vary with location, language, device, date, personalisation, authority and freshness. Different results do not always mean you need separate pages. Shared results do not guarantee that a single page is the best choice for your business either.
6. Make the SEO decision
The final cluster should help you make a practical decision. Depending on the project, it might represent:
one URL and one primary page brief;
several URLs within a shared topic area;
one content format, such as a guide, category page or comparison page;
one audience or business segment;
a site section or branch of your information architecture;
a group that needs manual review because the evidence is mixed.
This keeps the algorithm in a supporting role. What matters is not whether the groups look mathematically tidy, but whether they help you make reliable planning decisions.
How to evaluate a cluster
Clustering metrics such as silhouette score or Davies–Bouldin score can help check the shape of your results. These are intrinsic metrics: they assess the clustering itself, for example whether groups are compact or well separated under the chosen representation. They do not tell you whether those groups make useful SEO pages.
Include checks that focus on the decisions you need to make. For a sample of query pairs or proposed clusters, ask reviewers to record whether the queries have:
the same main task;
a compatible audience and stage of the customer journey;
a similar expected content format;
meaningful SERP overlap, measured under defined collection conditions;
the same likely target URL, or a clear reason to use separate pages;
important differences in product, location, brand or entity.
Check how consistently reviewers agree, too. If a query is ambiguous, record that rather than forcing everyone to treat one interpretation as the correct answer. Also distinguish between missing a paraphrase and wrongly merging separate needs. Both are errors, but their SEO consequences may differ.
Visual tools can help you inspect the dataset. Dimensionality-reduction methods such as PCA, t-SNE and UMAP turn high-dimensional representations into views that are easier to explore. A clear-looking map is not proof that the original clusters are stable, though. Its appearance may depend on the projection settings you choose.
What semantic clustering can and cannot demonstrate
Semantic clustering can plausibly improve discovery. It may find relationships that word matching misses and help you review a keyword list more systematically. For example, across 2,000 queries, it may expose connections between topics or highlight searches that do not fit your existing topic categories. These are practical possibilities suggested by how the method represents language, not established SEO performance results.
Finding useful relationships does not demonstrate that semantic clustering improves rankings, organic traffic, conversions or keyword cannibalisation. The evidence brought together here does not directly establish those outcomes as general benefits of semantic keyword clustering. That would require controlled research or carefully designed observational studies.
Likewise, statements about search engines understanding variations in language do not prove that a third-party embedding model works like a search engine. They do not establish that it reproduces the engine's internal representations, thresholds or ranking behaviour. Search-engine language matching and your team's clustering workflow are related subjects, but evidence about one cannot simply stand in for evidence about the other.
Limitations to address before operational use
Model dependence: switching embedding models can change similarity scores and which queries belong to each cluster.
Domain dependence: acronyms, product language and specialist terminology may be represented differently from everyday language.
Context loss: a short query may leave out the information needed to understand what someone means.
Threshold sensitivity: even small changes to a similarity threshold can wrongly merge queries or miss useful relationships.
SERP instability: search results change over time and across markets, devices and locations.
Intent ambiguity: a single query can have several reasonable interpretations, particularly when it is brief.
Operational mismatch: a mathematically coherent group may still be a poor fit for one page because of business needs, audience differences or content-format constraints.
These limitations do not make semantic similarity unusable. They help define where it belongs: early enough to improve discovery, but not treated as the final word on what content to create.
ClusterIQ Conclusion
Semantic similarity is most useful when you need to find relationships that matching words alone cannot reveal. Embeddings can surface paraphrases, synonyms and topic connections across a keyword dataset, making them useful for exploration and for suggesting candidate groups. How much they help depends on the model, the dataset and how you evaluate the results.
The key distinction is between semantic relatedness and shared SEO suitability. Similar language does not necessarily mean the same task, audience, search results, content format or URL. A strong workflow combines checks on exact wording with semantic candidates, intent analysis, SERP evidence and human judgement.
A useful cluster is more than a set of nearby vectors. It supports a decision you can explain about what to create, combine, separate or review. The next evidence gap is clear: compare lexical, semantic, SERP-only and hybrid methods on a labelled query set, then assess not just how neatly the queries group together, but how good and consistent the resulting SEO decisions are.

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

How to combine search intent, semantic similarity and SERP overlap for keyword clustering

Cosine similarity for keyword clustering: what the score really means
