Skip to main content
All articles
Clustering
24 September 2026 5 min read

Entity-aware keyword clustering: preserving brands, products and attributes

Semantic similarity can smooth over the exact entities that determine page ownership. Learn how brands, products, locations and attributes can become explicit clustering evidence.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Semantic embeddings are excellent at spotting broad meaning. That's genuinely useful, right up until the moment a small entity difference is the thing that actually matters.

"Nike Pegasus 41", "Nike Pegasus Trail 5" and "Nike Vaporfly 4" all sit in roughly the same semantic neighbourhood. For an ecommerce taxonomy, a comparison page or a content brief, treating those names as interchangeable would be a serious mistake.

Entity-aware clustering adds a second layer on top of semantic similarity. It doesn't just ask whether two queries are about the same kind of thing. It asks which specific people, brands, products, places, organisations or attributes the query actually refers to.

Why semantic similarity can smooth over important differences

Embedding models compress text into dense numerical representations. That's what lets them recognise paraphrases and related concepts, but the compression necessarily throws away some surface detail.

In general-language tasks, that's often exactly what you want. In SEO, a single modifier can define an entire page:

  • "Mira Sport shower" versus "Mira Advance shower";
  • "MacBook Air M5" versus "MacBook Pro M5";
  • "mortgage adviser Salisbury" versus "mortgage adviser Southampton";
  • "Salesforce CRM pricing" versus "HubSpot CRM pricing".

These queries can sit close together semantically while still needing completely different products, pages or locations.

What actually counts as an entity?

In natural-language processing, named entity recognition identifies spans of text belonging to categories like people, organisations, locations and similar defined types. SpaCy's entity recogniser, for example, tags spans of a document with labelled entities.

For SEO, the useful entity set is usually broader and more specific to your domain than a generic NER model provides out of the box. You may need to recognise:

  • brands;
  • product families;
  • model numbers;
  • services;
  • locations;
  • materials;
  • sizes;
  • standards or regulations;
  • software versions;
  • medical or technical terms.

Some of these are classic named entities. Others are structured attributes that function like entities for the purposes of the business decision.

Entity extraction and entity linking are different jobs

Spotting that "Apple" is an entity is not the same as working out which Apple is meant.

Entity linking connects a mention to a canonical identity. Google's Knowledge Graph Search API, for instance, exposes entity identifiers and types so different names or aliases can resolve to the same underlying entity.

That distinction matters when your dataset includes:

  • brand aliases;
  • abbreviations;
  • old and new product names;
  • businesses sharing the same name;
  • locations with ambiguous names.

A clustering system doesn't need a full global knowledge graph for every query, but it should know when a stable entity ID is more useful than raw string matching.

Entities can act as constraints as well as features

There are a few ways to bring entity information in. The simplest is to append recognised entity labels or canonical names onto the representation. Another is to build separate entity features and combine them with semantic similarity.

For the higher-risk distinctions, an entity can become an outright constraint:

  • don't automatically merge different product model IDs;
  • don't cross country or location boundaries without review;
  • keep regulated categories separate even when the language looks similar;
  • treat competing brands as related but distinct commercial targets.

This tends to be more reliable than just hoping the embedding model preserves every distinction consistently on its own.

A worked example: product queries

Take these queries:

  • "bosch series 6 dishwasher reviews";
  • "bosch series 6 dishwasher price";
  • "bosch series 8 dishwasher reviews";
  • "best bosch dishwasher".

A semantic model should correctly spot that all four are strongly related. An entity-aware representation can go further and expose:

  • brand: Bosch;
  • product category: dishwasher;
  • model family: Series 6 or Series 8;
  • task modifier: reviews, price or comparison.

That lets the system keep the broad topic visible while still keeping the model-family and task boundaries clear.

Entity information helps explain clusters

One weakness of dense embeddings is that nobody can look at an individual vector dimension and understand why a relationship exists. Entities add an interpretable layer on top.

If two queries end up grouped together, the interface can show:

  • shared entity: Bosch;
  • shared category: dishwasher;
  • different model: Series 6 / Series 8;
  • different modifier: review / price.

That's far more useful than just showing a similarity score of 0.87 with no explanation attached.

Don't assume generic NER is enough on its own

General NER models are trained around predefined labels and make assumptions about where entity boundaries sit. SpaCy's own documentation notes that its transition-based recogniser may not suit every span-identification problem.

SEO datasets often contain domain-specific strings that generic models simply miss:

  • SKU-like model numbers;
  • dimensions such as 1200x800;
  • abbreviated service names;
  • industry acronyms;
  • compound product names;
  • part codes.

For these, deterministic dictionaries, regex rules, catalogue data or site taxonomies can be more reliable than a general-purpose entity model.

Use the business's own data as an entity source

For ecommerce and large service sites, the strongest entity dictionary probably already exists somewhere inside the organisation:

  • product feed;
  • PIM;
  • CRM;
  • category taxonomy;
  • location database;
  • brand catalogue;
  • internal service list.

Matching keyword mentions against these controlled sources gives you a direct link between the language people search with and the site's actual information architecture. It's also much easier to audit than an entirely open-ended model.

Entity-aware graphs become more useful

Graph approaches can represent entities explicitly instead of folding everything into a single keyword-to-keyword similarity score.

A heterogeneous graph could contain:

  • keyword nodes;
  • entity nodes;
  • existing URL nodes;
  • category nodes.

Edges can then carry different meanings: "mentions", "semantically similar to", "currently ranks on" or "belongs to category". That's a lot richer than a graph where every edge simply means cosine similarity above some threshold. It also fits the broader argument in keyword graphs for SEO: relationships are more useful when their meaning stays explicit.

Entity overlap shouldn't become another magic score

It's tempting to write a rule like "same entity means same cluster". That's too crude.

"Apple Watch battery life", "Apple Watch battery replacement" and "Apple Watch battery not charging" all share the same entity but represent completely different tasks. Entity evidence should therefore work alongside:

  • semantic similarity;
  • intent;
  • page type;
  • SERP evidence;
  • business rules.

It's particularly valuable for stopping false merges and explaining why two semantically close queries should still stay apart.

A practical entity-aware workflow

  1. Define the entities that matter. Start from the business decision, not a generic NLP label set.
  2. Extract or match them. Combine NER with dictionaries and structured business data.
  3. Resolve aliases where possible. Keep stable IDs for canonical entities.
  4. Preserve attributes separately. Product model, location and task modifiers shouldn't disappear into one label.
  5. Generate semantic similarity. Use embeddings for broader language relationships.
  6. Apply constraints or weighting. Strengthen shared entities and protect critical entity differences.
  7. Expose the evidence. Show practitioners which entities influenced a grouping.
Practical principle: embeddings tell you that queries are related. Entities help tell you what, specifically, they're related through.

ClusterIQ Conclusion

Entity-aware keyword clustering isn't a replacement for semantic embeddings. It's a way of stopping broad semantic similarity from erasing distinctions the site actually cares about.

Brands, products, locations and technical attributes often determine page ownership more directly than general topical closeness ever will. Extracting those signals explicitly also makes clustering far easier to explain and debug.

The strongest approach mixes both: semantic representations for breadth, entity evidence for precision, and practitioner rules for the distinctions the business simply can't afford to lose.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.