Numbers and model codes in keyword clustering: why one digit can change the page
Dimensions, years, model numbers and capacities can carry more SEO meaning than the surrounding words. Learn how ClusterIQ protects numeric distinctions in semantic clustering.

Farky Rafiq
Founder of ClusterIQ

Numbers are often treated as background noise in data processing, yet in search, they frequently define the entire product. A "600mm vanity unit" and a "1200mm vanity unit" share almost every character, but they represent different physical needs and different inventory. Similarly, "iPhone 15" and "iPhone 16" are semantically identical to a basic algorithm, but the user intent is distinct. Even a year, like "2025 tax rates" versus "2026 tax rates", can demand entirely different content assets.
For ClusterIQ to be effective, it must distinguish between a numeric token that is just a modifier and one that changes the entity, the required page, or the temporal context.
Numbers have several SEO roles
When you are looking at a few thousand keywords from Ahrefs or Search Console, you will notice numbers appearing in various roles:
- dimensions and capacities;
- model and version numbers;
- years and dates;
- prices and quantities;
- technical specifications.
Treating all these as the same type of data leads to messy clusters. A width measurement requires different logic than a software version number.
Embeddings can overemphasise the shared words
Modern SEO often relies on vector embeddings to group keywords. The problem is that two product queries might differ by only one digit and remain extremely close in a mathematical sense. While this is helpful for broad topical grouping, it is dangerous for page mapping. If that single digit identifies a different product model, merging them onto one URL could ruin your content plan.
Extract typed numeric features
Rather than just seeing "600" as a string of text, ClusterIQ can identify the context. It looks for the value (600), the unit (mm), the attribute (width), and the entity context (vanity unit). By breaking the data down this way, a conflict between a 600mm and 800mm query becomes explainable and manageable rather than a random grouping error.
Worked example: boiler model codes
Consider a list of boiler queries:
- ABC 24 boiler;
- ABC 30 boiler;
- ABC 35 boiler.
These all belong to the same product range, but they correspond to specific output capacities and distinct product pages. ClusterIQ can maintain the relationship between the range while preventing the system from automatically merging these into a single page-level cluster based on the model code.
Normalise formatting, not identity
Users search inconsistently. "ABC-30", "ABC30", and "ABC 30" usually refer to the same thing. We use rules or catalogue aliases to canonicalise this formatting. This ensures that "ABC 30" stays grouped together regardless of hyphens, while remaining strictly separate from "ABC 35".
Units need canonical conversion carefully
In ecommerce, "60 cm" and "600 mm" are the same physical dimension. When handling product attributes, converting these to a canonical value helps group intent correctly. However, it is vital to keep the raw form and the converted value separate so the system can explain why it matched them in the final report.
Years can be temporal entities
A year in a query is rarely just a number. It often signals current regulations, annual reports, vehicle models, or specific events. We avoid removing years as generic stopwords because they are often the primary intent signal for the user.
Price numbers are different again
Queries like "Laptops under £500" and "laptops under £1,000" share a category but represent different budget segments. Depending on your site structure, the right treatment might be a filtered category page or a specific gift guide. ClusterIQ preserves these numeric constraints, allowing you to make the final URL decision in your ecommerce workflow.
Model-code patterns can be learned from catalogues
If you have access to product feeds with SKUs and model names, use them. These controlled values are far more reliable for identifying codes within search queries than relying on generic patterns or regular expressions alone.
Character n-grams can support formatting variants
Character n-grams are excellent for spotting near-identical codes that differ only by a space or a dash. However, we always suggest validating these matches against known catalogue data to ensure accuracy.
Numeric conflicts can become hard rules
If two queries map to mutually exclusive products, a high semantic similarity score should not be allowed to merge them. ClusterIQ can be configured to respect the shared parent topic while blocking a page-level merge that would confuse your site architecture.
Do not make every numeric difference a new cluster
Context is everything. "Top 5 SEO tools" and "top 10 SEO tools" usually serve the same editorial purpose. In this case, the number is a stylistic choice rather than a product differentiator. The meaning of the number always depends on the intent of the page.
Use distributions to describe demand
Instead of forcing every number into a new cluster, you can use them to describe a category. A report might show that 40% of demand is for 600mm units, while 30% is for 800mm. This helps with inventory and facet planning without cluttering your keyword map.
Version numeric rules
When you change how you handle unit conversions or year modifiers, it can shift how thousands of keywords are grouped. We treat these processing changes as distinct versions. This ensures that if your data changes, you know whether it was due to a shift in the market or a change in your settings.
Practitioner principle: one digit can be a minor modifier or the entire product identity. Preserve the numeric meaning before semantic similarity smooths it away.
ClusterIQ Conclusion
Numbers and model codes are vital pieces of evidence for technical and ecommerce SEO. By normalising formatting and protecting product identity, ClusterIQ helps you manage complex datasets without losing the specific details that drive conversions.
Data provenance matters as much as the method
It is essential to know exactly where your analysis came from. This means keeping track of the source dataset, the language, and the specific processing rules used at the time. Without this provenance, you might mistake a change in your internal logic for a genuine shift in how people are searching.
This is especially important when working at scale. Your data might include Search Console clicks, Semrush estimates, and manual business rules all at once. ClusterIQ keeps these sources attributable. When you refresh your analysis, you can see what changed in the raw data versus what changed in the model. This makes your reporting more reliable and prevents technical updates from accidentally rewriting your previous SEO insights.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Product attribute extraction for SEO: turning keyword modifiers into structured ecommerce evidence
Entity-aware keyword clustering: preserving brands, products and attributes
