Skip to main content
All articles
Clustering
20 July 2026 4 min read

Product attribute extraction for SEO: turning keyword modifiers into structured ecommerce evidence

Sizes, materials, colours, models and styles often determine ecommerce page ownership. Learn how to extract product attributes and combine them with semantic clustering.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Diagram showing ecommerce keyword phrases being converted into structured attributes such as size and mounting style, then grouped into separate product families.

When you are looking at thousands of ecommerce keywords from Ahrefs or Search Console, you quickly realise that semantic similarity only tells half the story. A query like "800mm matt black fluted shower screen" contains a wealth of specific data points. These modifiers are not just fluff; they are the exact attributes that dictate whether a keyword belongs on a category page, a specific facet, or a unique product URL.

ClusterIQ makes ecommerce clustering more precise by identifying these attributes explicitly. By extracting these details before the embedding model organises the broader meaning, we ensure the final groups reflect how your business actually organises its inventory.

What counts as a product attribute?

In a practical SEO context, attributes are the building blocks of your product descriptions. Common families include:

  • brand and manufacturer;
  • model name or specific range;
  • dimensions and size;
  • colour and finish;
  • material;
  • style or aesthetic;
  • fit, compatibility, or socket type;
  • capacity;
  • technical specifications;
  • target audience or specific use case.

The most useful schema for your project will always come from your actual catalogue, rather than a generic NLP textbook.

Use the product feed as a source of truth

If you are working with a client who has a PIM (Product Information Management) system or a robust product feed, you already have a "source of truth". These controlled values are gold mines for extraction.

For instance, a bathroom retailer’s catalogue might define a specific product as:

  • width: 800mm;
  • finish: matt black;
  • glass type: fluted;
  • door type: hinged;
  • brand: Merlyn.

By matching keyword language to these specific fields, ClusterIQ’s analysis becomes much easier to explain to stakeholders. It keeps your SEO terminology tied directly to real, shippable inventory.

Deterministic extraction is often better than a model

While AI is impressive, sometimes a simple rule is more reliable. Dimensions, SKUs, and known brand names can usually be matched using dictionaries and regular expressions.

A regex pattern can easily spot common dimension formats like "600x900mm". A brand dictionary ensures that "Nike" is always identified as the manufacturer, protecting exact product families from being blurred. We use semantic classification when a phrase requires genuine interpretation, but we stick to deterministic lookups where the data is clear and fixed.

Aliases need canonical values

Searchers are rarely consistent. In a list of three thousand keywords, you will see variations like:

  • black;
  • matt black;
  • matte black;
  • black finish.

While we preserve the raw wording for your content briefs, the structured attribute maps all these variations to a single canonical value from your catalogue. This helps ClusterIQ compare search demand against your actual stock levels without needing to rewrite the original query data.

Worked example: vanity units

Imagine you are sorting through a cluster of vanity unit keywords:

  • "600mm wall hung vanity unit oak";
  • "600 wall mounted vanity oak";
  • "800mm wall hung vanity unit oak";
  • "600mm floor standing vanity unit oak".

Standard semantic similarity might lump these all together under "oak vanity units". However, structured extraction preserves the vital differences:

  • width: 600 vs 800;
  • mounting: wall-hung vs floor-standing;
  • finish: oak.

These distinctions are what help you decide if you need one category page or three separate faceted URLs.

Attribute agreement can strengthen edges

When building a keyword map, shared attributes act as a signal of intent. In a weighted keyword graph, if two keywords share a product type, a width, and a mounting style, our confidence that they belong together increases significantly. Conversely, if the product models conflict, the relationship weakens.

Attribute conflicts can become hard constraints

Some differences are too significant to be ignored. If two queries refer to incompatible product models or sizes that require entirely different shipping logic or catalogue pages, ClusterIQ can flag this as a conflict. This prevents the system from automatically merging keywords that technically require separate landing pages.

Attributes support facet decisions

Once your keyword demand is mapped to catalogue attributes, you can make much smarter faceted-navigation decisions. You can compare search volume against your inventory depth and existing category coverage to see exactly where a new facet or filter is justified.

Attribute distributions describe the cluster

Instead of a vague cluster label, you get a data-driven profile. A group might be 80% "black finish" and 60% "800mm". This makes the resulting content plan or developer brief far more specific and actionable than a generic topic name.

Do not force every attribute into the cluster definition

Not every modifier deserves its own page. Colour might be a vital clustering factor for clothing but merely a filter for power tools. Use your business data and existing page performance to decide which attributes should actually constrain your grouping logic.

Keep extraction quality measurable

To ensure accuracy, we recommend testing a sample of your data. Look for precision, missed values, and how well aliases are being resolved. When you manually correct an attribute within ClusterIQ, that feedback helps refine the extraction rules for the rest of your dataset.

Product data improves explainability

It is much easier to explain to a client that two keywords were separated because their "mounting style" conflicted than it is to talk about vector space and embedding scores. This transparency makes the tool’s output far more trustworthy for experienced SEOs.

Practitioner principle: ecommerce queries are not only sentences. They are often compact product records written in natural language.

ClusterIQ Conclusion

Product attribute extraction adds a necessary layer of ecommerce intelligence to keyword clustering. By combining semantic meaning with catalogue-aware data like size, brand, and material, ClusterIQ creates a topic structure that respects the practical product distinctions that define your site architecture.

Use attribute extraction to improve data quality upstream

Sometimes, the clustering process highlights gaps in your own data. If ClusterIQ consistently finds user language that doesn't map to your existing attributes, it might suggest your product feed is missing key information. Fixing this doesn't just help SEO; it improves your site's internal search, filtering, and overall user experience. The clustering workflow becomes a stress test for how well your structured data matches the way customers actually speak.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.