Entity-aware keyword clustering: preserving brands, products and attributes
Semantic similarity can smooth over the exact entities that determine page ownership. Learn how brands, products, locations and attributes can become explicit clustering evidence.

Farky Rafiq
Founder of ClusterIQ
Semantic embeddings are excellent at spotting broad meaning. That's genuinely useful, right up until the moment a small entity difference is the thing that actually matters.
"Nike Pegasus 41", "Nike Pegasus Trail 5" and "Nike Vaporfly 4" all sit in roughly the same semantic neighbourhood. For an ecommerce taxonomy, a comparison page or a content brief, treating those names as interchangeable would be a serious mistake.
Entity-aware clustering adds a second layer on top of semantic similarity. It doesn't just ask whether two queries are about the same kind of thing. It asks which specific people, brands, products, places, organisations or attributes the query actually refers to.
Why semantic similarity can smooth over important differences
Embedding models compress text into dense numerical representations. That's what lets them recognise paraphrases and related concepts, but the compression necessarily throws away some surface detail.
In general-language tasks, that's often exactly what you want. In SEO, a single modifier can define an entire page:
- "Mira Sport shower" versus "Mira Advance shower";
- "MacBook Air M5" versus "MacBook Pro M5";
- "mortgage adviser Salisbury" versus "mortgage adviser Southampton";
- "Salesforce CRM pricing" versus "HubSpot CRM pricing".
These queries can sit close together semantically while still needing completely different products, pages or locations.
What actually counts as an entity?
In natural-language processing, named entity recognition identifies spans of text belonging to categories like people, organisations, locations and similar defined types. SpaCy's entity recogniser, for example, tags spans of a document with labelled entities.
For SEO, the useful entity set is usually broader and more specific to your domain than a generic NER model provides out of the box. You may need to recognise:
- brands;
- product families;
- model numbers;
- services;
- locations;
- materials;
- sizes;
- standards or regulations;
- software versions;
- medical or technical terms.
Some of these are classic named entities. Others are structured attributes that function like entities for the purposes of the business decision.
Entity extraction and entity linking are different jobs
Spotting that "Apple" is an entity is not the same as working out which Apple is meant.
Entity linking connects a mention to a canonical identity. Google's Knowledge Graph Search API, for instance, exposes entity identifiers and types so different names or aliases can resolve to the same underlying entity.
That distinction matters when your dataset includes:
- brand aliases;
- abbreviations;
- old and new product names;
- businesses sharing the same name;
- locations with ambiguous names.
A clustering system doesn't need a full global knowledge graph for every query, but it should know when a stable entity ID is more useful than raw string matching.
Entities can act as constraints as well as features
There are a few ways to bring entity information in. The simplest is to append recognised entity labels or canonical names onto the representation. Another is to build separate entity features and combine them with semantic similarity.
For the higher-risk distinctions, an entity can become an outright constraint:
- don't automatically merge different product model IDs;
- don't cross country or location boundaries without review;
- keep regulated categories separate even when the language looks similar;
- treat competing brands as related but distinct commercial targets.
This tends to be more reliable than just hoping the embedding model preserves every distinction consistently on its own.
A worked example: product queries
Take these queries:
- "bosch series 6 dishwasher reviews";
- "bosch series 6 dishwasher price";
- "bosch series 8 dishwasher reviews";
- "best bosch dishwasher".
A semantic model should correctly spot that all four are strongly related. An entity-aware representation can go further and expose:
- brand: Bosch;
- product category: dishwasher;
- model family: Series 6 or Series 8;
- task modifier: reviews, price or comparison.
That lets the system keep the broad topic visible while still keeping the model-family and task boundaries clear.
Entity information helps explain clusters
One weakness of dense embeddings is that nobody can look at an individual vector dimension and understand why a relationship exists. Entities add an interpretable layer on top.
If two queries end up grouped together, the interface can show:
- shared entity: Bosch;
- shared category: dishwasher;
- different model: Series 6 / Series 8;
- different modifier: review / price.
That's far more useful than just showing a similarity score of 0.87 with no explanation attached.
Don't assume generic NER is enough on its own
General NER models are trained around predefined labels and make assumptions about where entity boundaries sit. SpaCy's own documentation notes that its transition-based recogniser may not suit every span-identification problem.
SEO datasets often contain domain-specific strings that generic models simply miss:
- SKU-like model numbers;
- dimensions such as 1200x800;
- abbreviated service names;
- industry acronyms;
- compound product names;
- part codes.
For these, deterministic dictionaries, regex rules, catalogue data or site taxonomies can be more reliable than a general-purpose entity model.
Use the business's own data as an entity source
For ecommerce and large service sites, the strongest entity dictionary probably already exists somewhere inside the organisation:
- product feed;
- PIM;
- CRM;
- category taxonomy;
- location database;
- brand catalogue;
- internal service list.
Matching keyword mentions against these controlled sources gives you a direct link between the language people search with and the site's actual information architecture. It's also much easier to audit than an entirely open-ended model.
Entity-aware graphs become more useful
Graph approaches can represent entities explicitly instead of folding everything into a single keyword-to-keyword similarity score.
A heterogeneous graph could contain:
- keyword nodes;
- entity nodes;
- existing URL nodes;
- category nodes.
Edges can then carry different meanings: "mentions", "semantically similar to", "currently ranks on" or "belongs to category". That's a lot richer than a graph where every edge simply means cosine similarity above some threshold. It also fits the broader argument in keyword graphs for SEO: relationships are more useful when their meaning stays explicit.
Entity overlap shouldn't become another magic score
It's tempting to write a rule like "same entity means same cluster". That's too crude.
"Apple Watch battery life", "Apple Watch battery replacement" and "Apple Watch battery not charging" all share the same entity but represent completely different tasks. Entity evidence should therefore work alongside:
- semantic similarity;
- intent;
- page type;
- SERP evidence;
- business rules.
It's particularly valuable for stopping false merges and explaining why two semantically close queries should still stay apart.
A practical entity-aware workflow
- Define the entities that matter. Start from the business decision, not a generic NLP label set.
- Extract or match them. Combine NER with dictionaries and structured business data.
- Resolve aliases where possible. Keep stable IDs for canonical entities.
- Preserve attributes separately. Product model, location and task modifiers shouldn't disappear into one label.
- Generate semantic similarity. Use embeddings for broader language relationships.
- Apply constraints or weighting. Strengthen shared entities and protect critical entity differences.
- Expose the evidence. Show practitioners which entities influenced a grouping.
Practical principle: embeddings tell you that queries are related. Entities help tell you what, specifically, they're related through.
ClusterIQ Conclusion
Entity-aware keyword clustering isn't a replacement for semantic embeddings. It's a way of stopping broad semantic similarity from erasing distinctions the site actually cares about.
Brands, products, locations and technical attributes often determine page ownership more directly than general topical closeness ever will. Extracting those signals explicitly also makes clustering far easier to explain and debug.
The strongest approach mixes both: semantic representations for breadth, entity evidence for precision, and practitioner rules for the distinctions the business simply can't afford to lose.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Keyword preprocessing before clustering: clean the data without erasing intent

TF-IDF vs semantic embeddings for keyword clustering
