Mapping keyword clusters to existing URLs with embeddings
Embeddings can compare keyword clusters with existing page content at scale. Learn how to generate candidate URL matches without letting semantic similarity make the final mapping decision.

Farky Rafiq
Founder of ClusterIQ

Once you've clustered your keywords, it's tempting to start creating new pages straight away.
But on a site that already has plenty of content, the better first question is: which existing pages already cover these clusters?
Embeddings are useful here because they let you represent both your query clusters and your page content in the same kind of semantic space, so you can compare them at scale.
First, decide how to represent a cluster
You can represent a cluster in a few different ways:
- embed the cluster label;
- embed a concatenation of representative queries;
- average the embeddings of the member queries;
- use a weighted centroid that gives more weight to core or high-value terms.
The label on its own tends to lose too much detail. Representative queries or a centroid keep more of the underlying language intact.
Then represent each page properly
A page embedding shouldn't necessarily be built from the entire raw HTML.
Useful text to include is usually:
- the title;
- the H1;
- the main body copy;
- prominent headings;
- structured product or service descriptors.
Strip out navigation, boilerplate and repeated footer text, because that content makes unrelated pages look artificially similar to each other.
Retrieve several candidates, not just one winner
For each cluster, pull back the top few semantically closest URLs rather than a single answer.
That gives you a review set containing:
- candidate 1;
- candidate 2;
- candidate 3;
- the similarity evidence;
- any existing ranking evidence.
Resist the urge to automatically map the cluster to whichever page comes out top.
Semantic fit and page suitability aren't the same thing
A blog article can be semantically close to a commercial cluster simply because it discusses the same products.
That doesn't make it the right landing page.
Add page-type evidence to the mix, such as:
- category;
- product;
- guide;
- pricing;
- location;
- documentation.
A strong candidate needs to fit both the subject and the job the page has to do.
Search Console can confirm who already owns a query
If your candidate page already gets impressions for representative cluster queries, that's a good sign your mapping is right.
If a semantically close page has never shown up while another URL keeps appearing instead, dig into why before you change the target.
The query-page matrix workflow is a good source of first-party evidence for this step.
Use entities to catch false matches
Pages for similar products can end up with very close embeddings even when they shouldn't be treated as interchangeable.
Entity checks help protect distinctions that really matter, such as:
- brand;
- model;
- location;
- service type;
- audience.
Let semantic similarity generate the candidate set, then let entity compatibility filter it down.
Watch out for content length skewing the results
Long pages touch on more concepts, so they can end up looking broadly similar to lots of different clusters.
Try embedding focused page summaries or selected sections instead of treating every sentence as equally representative.
For big pages, section-level embeddings can show you exactly which part of the page matches the cluster.
Get a clean, canonical inventory of URLs first
Before you start matching, deduplicate:
- redirects;
- canonical duplicates;
- parameter variants;
- staging or non-indexable pages.
Mapping clusters to URLs that shouldn't even be indexed just creates confusion further down the line.
Use thresholds as review controls, not final answers
Rather than one fixed cut-off, try banding results:
- strong match: likely an existing-page candidate;
- moderate match: needs review;
- weak match: potential gap.
Calibrate these bands against page mappings you already know are correct.
A practical URL-matching workflow
- Create clean page text representations.
- Embed pages and cluster representatives using the same model.
- Retrieve top candidate URLs.
- Check page-type and entity compatibility.
- Add Search Console evidence.
- Classify as keep, improve, consolidate or create.
- Store confidence scores and any reviewer overrides.
Practitioner principle: embeddings are excellent at surfacing plausible existing pages. They're not a licence to skip the decision about what a page is actually for.
ClusterIQ Conclusion
Embedding-based URL matching can turn a slow, manual mapping exercise into a manageable candidate-review workflow.
Represent clusters and pages consistently, pull back several candidates, then layer in page type, entities and Search Console evidence. That way you keep the speed of semantic search while still applying the judgement that good information architecture needs.
Worked example: semantically close, but operationally wrong
A cluster about "enterprise SEO software" might look highly similar to a long educational article that explains enterprise SEO in general terms. But if the cluster is full of pricing, demo and platform queries, that article is a poor primary target, no matter how strong the semantic match looks.
Page-type filtering can remove that guide from the commercial candidate set while still keeping it around as a useful internal-link target.
Use section-level matching for broad pages
Large guides and category pages often cover several topics at once. Embedding the whole page as one block can blur those relationships. Splitting the main content into meaningful sections and retrieving the strongest section match shows you why a page was selected, and whether the cluster is already covered in enough depth.
This is particularly handy when you're deciding between "improve the existing page" and "create a new page".
Build a known-correct mapping set
Before rolling embedding matches out across the whole site, collect a set of clusters paired with pages that practitioners already agree are the right targets. Use those examples to check how often the correct URL shows up in the top one, top three and top five candidates.
That gives the retrieval layer a clear, measurable job: surface the right page for review. It doesn't need to make the final call on its own.
Store rejected candidates too
If a reviewer rejects a highly similar page because the page type, entity or commercial role is wrong, hang on to that decision. Those rejections are valuable evidence for improving your candidate filters and for building future QA benchmarks.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Keyword cannibalisation: how clustering helps find overlap without inventing a problem

Building a Search Console query-page matrix for cluster analysis
