Skip to main content
All articles
Clustering
1 June 2026 4 min read

Stemming, lemmatisation and embeddings in SEO: when normalising words helps and when it destroys meaning

Traditional text pipelines often reduce words to roots or lemmas. Learn when that still helps SEO clustering and when modern embeddings should see the original wording.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing raw query wording branching into an embeddings path, a selectively normalised lexical-features path, and a protected entity bypass.

If you have ever exported a few thousand keywords from Ahrefs or Search Console, you have likely noticed how messy the data is. You will see "running shoe", "running shoes", and "run shoe" all appearing as separate rows. In traditional Natural Language Processing (NLP), the standard response is to normalise these words before doing any heavy lifting. We usually do this through stemming or lemmatisation.

Stemming is a bit like using a machete; it chops off the ends of words to find a common root. Lemmatisation is more like using a scalpel, mapping words back to their actual dictionary form based on how they are used. Both techniques aim to reduce variety so your data is easier to handle, but in the world of SEO, they can sometimes strip away the very nuances that define search intent. Modern tools like ClusterIQ have to be careful here: we only normalise when the benefit is clear, because modern embeddings are often smart enough to understand these relationships without us interfering.

What stemming does

A stemmer follows a set of rigid rules to strip prefixes and suffixes. It does not care if the result is a real word. For example, "optimising", "optimised", and "optimisation" might all be hacked down to "optim". While this makes your list of unique terms smaller and easier for a computer to process, it is a very blunt instrument.

What lemmatisation does

Lemmatisation is more sophisticated. It looks at a word and tries to find its "lemma" or base form. It understands that "running" should map to "run" if it is used as a verb, but it might treat it differently in another context. It is linguistically more accurate than stemming, though it requires more processing power to get right.

Why SEO data is risky

The problem we face as SEOs is that keyword strings are incredibly short and lack context. When you are looking at a list of 2,000 keywords, a word that looks like a simple variation might actually represent a completely different commercial intent. Product model codes, brand names, and specific locations often look like ordinary words to a basic algorithm. If we over-normalise, we risk merging distinct search intents before our clustering model even gets a look at them.

Worked example: singular and plural categories

Take the terms "running shoe" and "running shoes". If you are using a modern embedding model, these are usually seen as related enough that you do not need to stem them. However, if you are using a TF-IDF approach to group your keywords, lemmatisation can be helpful to reduce "noise" and show that these two terms are essentially the same thing. The right choice depends entirely on how you are representing the data, not a blanket rule that plurals are always bad.

Exact query preservation is essential

Even when we use normalisation, ClusterIQ keeps the original, raw query. This is vital for your workflow. It means when you are building a content plan or a page brief, you can see exactly what people typed into Google. It also allows you to audit the data to make sure a "smart" preprocessing step hasn't accidentally merged two things that should have stayed separate.

Use different preprocessing for different feature layers

A sophisticated way to handle this is to keep different versions of the text for different tasks. You might use:

  • The raw text for your embeddings and final reports;
  • A lightly cleaned version to find exact duplicates;
  • A lemmatised version for specific lexical analysis;
  • Separate lists for brands and locations.

This ensures that one transformation doesn't ruin the data for every other part of your SEO strategy.

Embeddings usually benefit from natural wording

The AI models we use today were trained on natural language, like books and websites. If you feed them chopped-up text like "optim search cluster", you are giving them something they don't recognise. It is often better to leave the keywords exactly as they are so the model can use its full understanding of English to group them correctly.

Lemmatisation can still help sparse retrieval

If you are using older methods like TF-IDF, lemmatisation is still a great ally. It prevents your data from becoming too fragmented by grouping different forms of the same word together. You can read more about this in ClusterIQ's TF-IDF versus embeddings comparison to see which approach fits your specific project.

Language matters

It is important to remember that the rules for English do not work for German, Spanish, or French. If you are working on a multilingual site, your preprocessing needs to be specific to that language or kept to an absolute minimum to avoid creating a mess of your international keyword research.

Do not stem entities blindly

Brand names and specific product ranges should be protected. You don't want a tool to "correct" a brand name into a common noun. We try to identify these entities first so they can be left alone while the rest of the text is processed.

Measure downstream effects

Before deciding on a method, we look at the results. Does lemmatisation actually make the clusters better? Does it help find more duplicates? If adding a complex step like lemmatisation doesn't actually change the final decisions you make for your URLs or content briefs, it is usually better to keep things simple.

Use preprocessing versioning

Any time you change how you process your keywords, it can change your final clusters and reports. We treat these changes with the same respect as a model update, recording them in the ClusterIQ manifest so you can always trace back why your data looks the way it does.

Practitioner principle: normalise text to remove technical noise, not to erase linguistic distinctions before you know whether they matter.

ClusterIQ Conclusion

Stemming and lemmatisation are useful tools, but they aren't always necessary for modern semantic clustering. By keeping the raw language for embeddings and using normalisation only where it adds value, ClusterIQ ensures your SEO data remains accurate and actionable.

Failure modes are part of the feature, not an appendix

In a real-world SEO environment, knowing where a tool might fail is just as important as knowing where it succeeds. Whether it is a rare brand name, a very short ambiguous phrase, or a complex multilingual query, these edge cases often define the quality of your final content plan. Instead of forcing a "clean" but wrong answer, it is often better to show uncertainty. This might mean leaving a query unassigned or presenting a few different ways a keyword could be grouped, allowing you to make the final call.

By categorising these failures, we can figure out if the issue is in the initial cleaning of the data or the clustering itself. This allows for much more precise improvements to your workflow rather than just guessing why a certain keyword ended up in the wrong bucket.

Keep a transformation diff for difficult cases

When a keyword is changed by preprocessing, ClusterIQ can show you the "before and after". This is incredibly helpful when you see a keyword in a cluster where it doesn't seem to belong. You can quickly check if a lemmatiser accidentally stripped away a crucial bit of meaning.

We use a specific set of test words to monitor this, including brands that look like regular words and terms where a small change in spelling changes the entire meaning. By keeping our semantic models and our lexical models independent, we ensure that a shortcut taken in one area doesn't negatively impact your entire SEO analysis.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.