Weighting titles, headings and body text in page embeddings for SEO
A page title, H1 and body copy describe different aspects of page purpose. Learn how weighted or multi-field embeddings can improve URL matching without hard-coding SEO folklore.

Farky Rafiq
Founder of ClusterIQ

When you map 1,000 keywords from Ahrefs, Semrush or Search Console to existing URLs, a familiar problem appears: a page mentions a topic, but that does not mean it should own it.
Separating titles, headings and body copy can make those decisions easier to check, saving review time and producing clearer content plans. An embedding, a numerical representation of meaning, is useful here, but a page is not one undifferentiated block of text.
Titles and H1s usually describe its main purpose. Headings organise sections; body copy supplies detail and examples. Product grids, breadcrumbs and structured data add context. If everything is combined into one embedding, the longest field can overshadow shorter, more revealing evidence.
Why field structure matters
Consider a commercial category page with:
- title: “Black Shower Screens”;
- H1: “Black Shower Screens”;
- 1,500 words of buying guidance covering glass, sizing and installation.
A whole-body embedding may strongly associate it with installation and sizing because those subjects occupy most of the text. The title and H1 still describe what the page primarily is: a shopping category.
Do not assume title text is always superior
Titles can be outdated, over-optimised, templated, too generic or misaligned with visible content. Giving them automatic priority simply creates a different source of error.
For ClusterIQ, field evidence should be useful evidence, not a supposed universal SEO weighting that makes titles inherently more trustworthy.
Approach one: concatenate fields with labels
The simplest approach joins fields together with explicit labels:
Title: Black Shower Screens. H1: Black Shower Screens. Breadcrumb: Bathrooms > Showers. Body: …
This gives the embedding model structural cues while keeping implementation straightforward. It still produces one representation, rather than separate evidence for each field.
Approach two: create separate field embeddings
A more explainable design would encode the title, headings and body separately, then preserve individual scores:
- title similarity;
- heading similarity;
- body similarity;
- entity agreement.
Instead of seeing only one pooled score, a reviewer could see which evidence supports a proposed URL mapping and which evidence does not.
Approach three: weighted combination
A page-relevance model can combine normalised field scores, adjusted onto comparable scales, using explicit weights.
A mapping experiment might test whether title/H1 evidence deserves more influence over page-purpose matching than a peripheral body section. Learn or calibrate those weights against known page mappings. Do not choose 0.6/0.3/0.1 merely because the proportions look sensible.
Worked example: informational section on a commercial page
Suppose the category has a long FAQ about “how to clean shower screens”. A whole-page embedding retrieves it strongly for cleaning queries, while its title, H1 and product set indicate a commercial role.
A field-aware approach could recognise the FAQ as supporting relevance without automatically assigning a broad cleaning cluster to the category. In a content plan, that distinction helps you review whether the cluster needs its own informational page rather than expanding the category brief.
Headings can improve chunk context
When long pages are split into chunks, keep the heading path attached. A section labelled “Installation” becomes more informative when its page title and parent heading remain available.
That context is one reason section-aware chunking is stronger than cutting text into fixed token slices.
Breadcrumbs can add taxonomy context
Breadcrumbs show where a page sits in the site hierarchy. For ecommerce, they can help distinguish an identically named product from its parent category or identify the market context.
Keep this evidence separate: a poor taxonomy should not silently dominate the page representation.
Structured data can provide entity evidence
Product, Article or Organization structured data can supply useful page-type and entity signals. These fields may work better as structured features than as extra text poured into an embedding.
Template boilerplate needs control
Headers, footers and repeated navigation can introduce identical text across thousands of pages. Exclude or downweight this boilerplate before creating body embeddings so page-specific content drives the representation.
Use field disagreement as a quality signal
If a title says “Keyword Clustering Software” but the body embedding resembles a general SEO glossary, investigate the mismatch. It may indicate:
- thin commercial content;
- misleading metadata;
- template contamination;
- incorrect page classification.
The disagreement is a review prompt, not proof of any single problem.
Evaluate against real URL mappings
Build a benchmark of clusters with approved target URLs, then compare:
- body-only retrieval;
- title + body;
- separate field scores;
- chunk-aware retrieval;
- field + page-type evidence.
The winning representation is the one that improves the mappings that matter, not the one with the most elaborate configuration.
Do not turn weights into hidden policy
If ClusterIQ uses weighted combinations, the explanation should retain component scores. A reviewer should see, for example, that strong title and entity evidence outweighed moderate body similarity. That makes a recommendation easier to challenge before implementation.
Different page types may need different field strategies
Products, categories, articles and tools have different text structures. Validated, page-type-specific configurations may work better than one universal weighting.
Practitioner principle: page fields describe different aspects of purpose. Preserve those distinctions before compressing the page into one similarity score.
ClusterIQ Conclusion
Field-aware embeddings offer a way to make ClusterIQ's URL matching more reliable and explainable. Titles and headings should neither receive arbitrary authority nor disappear beneath long body copy. Use benchmarked mappings and transparent component evidence to find the balance.
How this fits the wider ClusterIQ workflow
This method belongs within a wider evidence chain: clean source data, preserve lexical and entity features, create semantic relationships, test the structure, then connect groups to pages, internal links or roadmap actions. Keeping stages separate makes recommendations easier to explain and safer to change.
The product should let an SEO challenge a result without becoming a data scientist. A concise explanation should show the strongest supporting evidence, the main conflict and the practical consequence of acceptance. Deeper diagnostics should retain parameters and scores.
After implementation, monitoring the same topic ID through Search Console, crawl data, inventory or other relevant sources would close the loop. The aim is to record whether approved actions improve ownership, coverage or usability over time, not merely describe today's structure.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Chunking long pages for SEO embeddings: preserving section relevance without fragmenting the page

From keyword clusters to URL mapping: turning groups into site architecture
