Near-duplicate page detection for SEO: combining shingles, embeddings and page structure
Near-duplicate pages can evade exact matching because wording changes while structure and purpose remain the same. Learn how layered similarity creates a stronger SEO review queue.

Farky Rafiq
Founder of ClusterIQ

When two pages look almost the same, should you merge them or keep both? Exact copies are relatively easy to spot. Near duplicates need more judgement.
Pages can reorder sections and add unique paragraphs while still doing the same job. Others share 60% of their words through a template but serve different products. ClusterIQ can combine several evidence layers into a review queue, helping you prioritise decisions rather than inspect every pair.
Start with technical duplicates
Normalise obvious URL and HTML variations first:
- protocol and host variants;
- tracking parameters;
- print versions;
- known canonical variants.
These do not need meaning-based analysis to diagnose. Separating them leaves a more useful set for content review.
Text shingles catch copied sequences
To spot copied wording despite different HTML, compare sets of adjacent words or characters, called shingles.
High overlap can reveal templated descriptions and lightly rewritten content. It tells you how much wording is shared, not whether both pages deserve to exist.
Embeddings catch meaning after rewriting
Pages can communicate nearly identical information without sharing exact phrases. Dense embeddings, numerical representations of meaning, help identify this kind of similarity.
The risk is mistaking related topics for duplicates. Always consider what each page is meant to do alongside its semantic similarity.
Worked example: local service pages
A business has 50 city pages where only the location name and two local sentences change.
Shingle overlap and embedding similarity are extremely high. Page type and section structure also match. ClusterIQ can flag this as a scaled near-duplicate pattern for review, rather than treating every URL as independent local content.
Structure similarity adds another layer
Alongside wording and meaning, compare:
- heading sequence;
- section count;
- template modules;
- word count by section;
- product or entity set.
Near-identical structure strengthens the case that pages may be duplicates, though it is not proof by itself.
Boilerplate should be removed or segmented
Headers, footers, legal copy and delivery modules can dominate wording comparisons. Remove these repeated components from the comparison or analyse them separately, so they do not overwhelm the main content.
Product-set overlap is decisive for category pages
Two categories with 95% identical products and very similar query clusters deserve closer review, even with different introductions.
Comparing product sets can expose this redundancy more directly than text embeddings: different copy may disguise almost the same range.
Query ownership closes the loop
Check which queries lead to each page. A Search Console export of a few hundred keywords can make this practical.
Shared query clusters strengthen the case for consolidation. If pages attract distinct demand, consider making their differences clearer rather than removing one.
Use several thresholds, not one magic score
A practical review rule can combine:
- high shingle overlap;
- high semantic similarity;
- same page type;
- high product/entity overlap;
- shared query ownership.
ClusterIQ can explain which conditions triggered the review, giving reviewers evidence to check rather than an unexplained score.
Cluster near-duplicate families
At scale, reviewing 10,000 page pairs becomes unwieldy. Build connected families instead: if Page A resembles B and B resembles C, review the family as one potential template or content problem.
Do not auto-delete the weaker page
Lower traffic does not necessarily mean a page should go. It may have better backlinks, a cleaner URL, stronger conversions or the right long-term role.
Use ClusterIQ's keep/improve/merge/remove framework to guide the decision.
Near duplicates can expose CMS problems
Look for recurring causes:
- faceted URL generation;
- location templates;
- duplicate category logic;
- content syndication;
- legacy migrations.
Fixing the template or generation logic may be more valuable than editing pages individually.
Track families after remediation
After consolidation or template changes, rerun detection and check that:
- the duplicate family has shrunk;
- canonical signals align;
- query ownership has consolidated;
- internal links point to the intended page.
Keep thresholds page-type specific
Product pages naturally share more template text than editorial guides. One universal wording-overlap threshold will produce many false positives.
Practitioner principle: use wording, meaning, structure, entities and search ownership together. Near-duplicate detection works best as layered evidence.
ClusterIQ Conclusion
Near-duplicate detection needs more than exact text matching. ClusterIQ can combine shingles, embeddings, page structure, product sets and Search Console ownership to prioritise page families for review, while avoiding false consolidation of legitimately related content.
Turn duplicate families into a remediation plan
Once ClusterIQ identifies a family, classify its members. One may be the intended canonical owner, another a legacy URL, another a legitimate local variant and another an accidental template duplicate. Reviewing them together makes remediation safer.
Create an action plan: nominate a survivor where appropriate, record unique content and query clusters to preserve, specify redirects or canonical actions, and list internal links to update. For ecommerce, keep product overlap and inventory differences visible so a genuinely distinct range is not lost.
Afterwards, ClusterIQ can rerun family comparisons and page-ownership checks. Success means clearer ownership with useful distinctions preserved, not simply fewer similar URLs.
Sample false positives deliberately
Test every rule against similar-looking pages that must remain separate. Product variants, international pages and templated support content make useful negative controls. They help prevent a more aggressive threshold from looking better merely because it flags more pages.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Mapping keyword clusters to existing URLs with embeddings

Keyword clustering for faceted navigation: deciding which filter combinations deserve pages
