How to combine search intent, semantic similarity and SERP overlap for keyword clustering
Search intent is a useful starting point for keyword clustering, but it cannot decide by itself whether terms belong on one page. This guide shows how to combine intent hypotheses, semantic similarity and contextual SERP overlap before assigning same-page, related-page or separate-content targets.

Farky Rafiq
Founder of ClusterIQ

You have 800 keywords in a spreadsheet and need to turn them into a sensible content plan. Which terms belong on one page, which need their own pages, and which are simply related topics?
A common starting point is to label each keyword by search intent: group the informational queries, separate the transactional ones and assign each group to a page. That is useful groundwork, but it can also lead to poor page decisions.
Two queries may appear to express the same intent yet need different pages. Others may use quite different wording but return many of the same search results. Comparing meaning helps you spot relationships between queries. Comparing search results gives you evidence about how the search engine currently responds to them.
So the most useful question is not just “what is the intent of this keyword?” It is:
Do these queries represent a sufficiently similar need, topic and search-result expectation to be served by the same page?
To answer that well, consider three signals:
- Intent hypotheses: what the person searching may be trying to achieve.
- Semantic similarity: how closely the queries relate in language and meaning.
- SERP overlap: how often the queries currently return the same documents or URLs on their search engine results pages.
None of these proves that two queries belong on the same page. Together, with human review of the less clear-cut cases, they give you a stronger basis for choosing between one shared page, related pages or separate content.
Search intent is a hypothesis, not a property you can read directly from a keyword
Search intent is a useful way to describe the goal behind a query. Common broad categories include navigational searches, where someone wants a particular website or page; informational searches, where they want information; and transactional searches, where they want to take an action. More detailed classifications cover goals such as learning, comparing, finding a specific product, completing a task or solving a problem.
These categories help you make sense of a keyword list and provide a useful first filter. For example, mirrorless camera may suggest broad product research, while buy Sony mirrorless camera is more clearly commercial.
But the words in a query do not tell you everything about the person typing them. You may not know their experience, urgency, audience, budget, preferred format or previous activity. Some searches are deliberately broad. Best mirrorless camera could come from a beginner, a professional photographer or someone replacing a compact camera.
For clustering, treat an intent label as a provisional hypothesis: a plausible explanation of what someone needs, rather than proof that every query with that label belongs on one page.
Intent may be enough for a tightly defined set of near-equivalent terms, such as searches for close product variants within one category. Relying on it becomes riskier when queries differ in audience, specificity, modifiers, content format or expected action.
Why similar intent can still lead to different pages
Imagine these terms appearing in a camera retailer’s keyword list:
- mirrorless cameras
- best mirrorless cameras
- mirrorless cameras for beginners
- mirrorless camera lenses
- how to use a mirrorless camera
You could label several of them informational or commercial investigation, meaning research before a possible purchase. But that label does not tell you which pages to build.
Mirrorless cameras may call for a category or buying page. Best mirrorless cameras may need an editorial comparison that explains its recommendations. Mirrorless cameras for beginners adds a specific audience, which could justify a dedicated guide or buying page. Mirrorless camera lenses shifts the focus to a different product, while how to use a mirrorless camera points towards practical instruction.
All five terms relate to the same broad topic, and some may share an intent category. Their likely page purposes still differ.
This is the distinction between topic relatedness and page equivalence: being connected to the same subject is not the same as needing the same page. A useful clustering process recognises both relationships.
What semantic similarity adds
People rarely use consistent wording when they search. Semantic similarity helps you find queries that are close in meaning, even when they do not share many words.
Sentence-embedding models do this by representing queries as vectors, or numerical representations, in a mathematical space. The system can then compare those vectors. A common measure, cosine similarity, compares the angle between them.
For example, a semantic model may recognise relationships between:
- project management software for small teams and team task management tool;
- cancel a streaming subscription and how to end my streaming membership;
- carry-on luggage dimensions and airline cabin bag size limits.
Matching words alone, often called lexical matching, may understate these connections. Embeddings can help you suggest possible groups, identify differently worded versions of a query and find wider topic groups that exact word matching would struggle to reveal.
There is no universal interpretation of a similarity score, though. What a score means depends on the model, its training data, the language, the subject area and the similarity measure used. A threshold that produces useful groups in one keyword list may work poorly in another.
Being close in meaning also does not guarantee identical search intent. Similar queries can differ in:
- audience or level of expertise;
- commercial stage;
- geographic or local requirement;
- product specification;
- desired content format;
- urgency or task complexity.
Embedding models can draw topical connections too broadly. They may link terms because they appear in similar language, even when the searches need different information or page types.
Use semantic similarity to generate candidates and expose relationships, not to assign queries to URLs without review.
What SERP overlap adds
SERP overlap measures how many documents or URLs two queries share within a defined set of search results. If two searches return eight of the same URLs in their top ten results, they have high overlap under that measurement.
Where semantic similarity compares meaning, SERP overlap asks:
How is the search engine’s retrieval system currently responding to these queries?
That makes it useful when deciding whether queries might share a page. If two searches return many of the same relevant pages under the same conditions, you have practical evidence that compatible content may serve them both.
Suppose carry-on luggage dimensions and cabin bag size limits produce substantially overlapping results in the same market and device context. That does not prove they need one URL, but it gives you a reason to investigate a combined page. One well-structured guide might address both needs.
Now suppose best mirrorless cameras returns comparison articles, while mirrorless camera lenses returns product categories and lens buying guides. Those results support separate page purposes, even though the terms belong to the same broad camera topic.
SERP evidence depends on context and can change. Results vary by location, language, device, date and result depth, meaning how far down the results you look. How you handle SERP features, changes in rankings and the search engine’s index coverage can also affect measured overlap. A single low-overlap snapshot might reflect collection conditions or fluctuating results, rather than a lasting difference in what pages people need.
You also need to inspect the shared results. The same URLs may appear because those pages cover a broad subject, not because every result addresses both needs. Check each page’s purpose, the products or concepts it covers, relevant query modifiers, its format and how it answers each search.
A staged workflow for multi-signal keyword clustering
A workable process keeps mathematical grouping separate from the final decision about which pages to create or use. For a list of a few hundred to a few thousand keywords, the following stages give you a way to organise the evidence before committing to a content plan.
1. Define the decision before clustering
Start with what you need the output to help you decide. Are you looking for:
- terms that can target the same URL;
- closely related pages within one site structure;
- separate content opportunities;
- or broad topic groups for further research?
These are different jobs. A cluster that works well as a research topic may be far too broad for one page. A group intended to target one URL usually needs a tighter relationship between its terms.
Keep cluster membership and page architecture as separate decisions. Membership describes a relationship in the data. Page architecture is an editorial and commercial choice about what a page should contain and who it should serve.
2. Normalise the keyword data
Before comparing queries, tidy up obvious inconsistencies. Remove duplicates, standardise letter case and preserve meaningful modifiers. Record useful fields such as language, market, device, search volume and where the keyword came from.
Be careful not to strip out words that change the need behind a search. Phrases such as for beginners, near me, pricing, template, support and download may look like small additions, but they can change the purpose of the page someone expects.
3. Create provisional intent labels
Assign broad intent labels, adding more detail where it helps. A keyword list for a SaaS business, for example, might distinguish between:
- category discovery;
- feature research;
- pricing and commercial evaluation;
- implementation guidance;
- templates or downloadable resources;
- support and troubleshooting.
Record how confident you are in each label instead of forcing every query into a definite category. Leave ambiguous terms visibly ambiguous. A low-confidence label tells the reviewer where a closer look may be worthwhile.
4. Generate semantic candidate groups
Use word-based rules, embeddings or a combination of both to identify related queries. Different clustering methods can help you explore those relationships:
- Hierarchical agglomerative clustering builds nested groups, which can be useful when a broad topic contains several tighter subgroups.
- Density-based approaches can identify irregularly shaped groups and possible outliers without forcing every term into a cluster.
- Graph-based methods represent relationships as connections, helping you spot communities, terms that bridge groups and weakly connected items.
Choose a method that fits the dataset and what you want to do with it. No single algorithm or threshold can turn every keyword list into the right site structure.
Keep the underlying scores and settings. An SEO specialist should be able to inspect, question and revise a model-produced cluster, rather than having to accept it as an unexplained result.
5. Validate promising pairs and groups with SERP evidence
For terms that look like possible same-page candidates, collect search results under documented conditions. At a minimum, record:
- market and location;
- language;
- device;
- collection date;
- result depth;
- URL normalisation rules;
- how duplicate URLs and SERP features are handled.
Look at both the amount of overlap and the types of pages appearing. A basic overlap measure can help, but ranking positions may matter too. Two queries that share only their number-one result present different evidence from two queries that share most of their top ten.
A chosen overlap threshold is not a universal rule. Its meaning depends on the query type, the result window you measured, the market and the quality of the pages returned.
6. Review the boundaries, not just the centres
The obvious members of a group usually need less attention than the uncertain ones. Focus human review where signals disagree, particularly terms with:
- high semantic similarity but low SERP overlap;
- low semantic similarity but high SERP overlap;
- conflicting intent labels;
- audience, location or format modifiers;
- weak or unstable cluster membership;
- large differences in commercial value or business purpose.
For each borderline case, ask what a page would actually need to contain to satisfy the query. Would it have to address distinct audiences, actions or formats without a clear purpose tying them together? If so, separate pages may be easier to justify.
If the differences fit naturally into sections, filters or supporting content, a shared page may be appropriate. The aim is to make a useful page, not just to keep a group of keywords together.
7. Assign page architecture after the evidence review
Use three practical outcomes:
- Same-page group: the queries describe a sufficiently similar need, and their search results and content requirements are compatible.
- Related-page group: the queries belong to one topic but need distinct pages, connected through the site structure and internal links.
- Separate-content group: the terms differ substantially in audience, task, product, format or observed search-result behaviour.
This three-way choice is often more useful than a binary “cluster” or “do not cluster” decision. Pages can support the same wider strategy without doing exactly the same job.
How to evaluate the clustering itself
Technical evaluation helps you check whether the groups are coherent under the way you have represented and compared the queries. Silhouette analysis, for example, compares how closely each item relates to its own cluster with how closely it relates to neighbouring clusters. A higher score may indicate that the groups are more cohesive internally and better separated from one another under those conditions.
That is not the same as being useful for SEO. A mathematically tight cluster can still combine queries that need different page formats. Equally, a useful same-page group may include borderline terms and receive a less tidy score.
If you have a reference set of human judgements, the Adjusted Rand Index can compare two sets of cluster assignments while adjusting for agreement that could happen by chance. It needs reference labels, however. SEO judgements may also be subjective, hierarchical or overlapping, rather than one objectively correct way to divide every term into a group.
Treat these metrics as diagnostic tools, not final proof that a group belongs on one page. Also track practical measures: how often reviewers agree, how many terms need manual intervention and how stable the clusters remain when the dataset or SERP collection date changes.
Common failure modes
Forcing every query into one intent category
Broad labels can hide differences that matter to the page decision. Add confidence levels and secondary attributes rather than treating every query as though it has only one possible goal.
Using an embedding threshold as a page rule
A cosine-similarity threshold depends on the model and dataset. It can flag terms worth comparing, but it cannot define a universal boundary for which queries belong on the same page.
Treating one SERP snapshot as permanent evidence
Record the collection conditions. Where a decision matters, check whether the results remain stable across multiple dates or environments. Today’s SERP shows current retrieval behaviour, not a permanent definition of what users want.
Assuming high overlap means identical content
Inspect the shared pages rather than relying on the overlap count. A broad domain may rank for several neighbouring queries while providing different sections or experiences for each need.
Optimising for cluster neatness
Tidy groups are satisfying, but your site structure needs to serve users, business objectives and the content’s actual scope. A slightly untidy group may work better in practice than a mathematically pure one.
What this method can and cannot tell you
Using several signals reduces the risk of making page decisions from one incomplete view of a keyword. Intent helps you describe the likely need. Semantic similarity finds related language. SERP overlap shows how the search engine currently groups results. Human review connects that evidence to what a page should do.
The process still cannot prove that one URL will rank for every term, that separate pages will perform better or that a cluster represents one stable user goal. Search results may deliberately offer variety to satisfy several interpretations of a query. Different site structures may also be reasonable for different markets, audiences and resource constraints.
The strongest claim you can reasonably make is a modest one: combining complementary signals gives you a better basis for investigation than relying on intent, embeddings or SERP overlap alone. How much weight to give each signal remains unresolved.
ClusterIQ Conclusion
Search intent remains an essential part of keyword clustering, but it is a working interpretation of someone’s need, not a complete answer to the page question. Semantic similarity can reveal connections that word matching misses. SERP overlap adds observed evidence about how searches are being served today.
Keep the distinction clear: intent is a hypothesis, while SERP overlap is an observation. Neither proves that queries belong on the same page. Use both alongside semantic evidence, document how you collected the results and inspect cases where the signals disagree. Make the page-architecture decision after that review, not before it.
There is still a useful research question for SEO teams: which combination of intent labels, lexical features, embeddings and SERP overlap best predicts expert agreement on same-page, related-page and separate-content decisions? Until that is tested on a documented dataset, this staged workflow is a sensible way to organise the evidence, not a proven universal formula.

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

When semantic similarity helps keyword clustering, and when it does not

Keyword clustering vs topic clustering: what’s the difference?
