Skip to main content
All articles
Clustering
24 September 2026 7 min read

Keyword preprocessing before clustering: clean the data without erasing intent

Keyword cleaning can improve clustering or quietly destroy useful distinctions. Learn which transformations are safe, which need testing and why raw queries should always be preserved.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing equivalent keyword variants converging through safe normalisation while meaningful product differences remain separate.

Keyword clustering actually starts before you ever run the clustering algorithm.

Feed a model noisy, duplicated or inconsistently encoded data, and even a strong embedding model ends up spending its effort modelling formatting quirks instead of real relationships. But swing too far the other way with aggressive cleaning, and you can just as easily erase the distinction between two products, two intents or two genuinely different pages.

So the preprocessing layer really has one job: strip out variation that's operationally meaningless while keeping variation that could change the SEO decision.

Why preprocessing matters more than it looks

A raw keyword export usually contains several kinds of noise at once:

  • leading and trailing whitespace;
  • different Unicode representations of the same visible text;
  • upper- and lower-case variants;
  • punctuation differences;
  • duplicate rows from merged data sources;
  • tracking columns and market labels;
  • spelling variants;
  • near-duplicates created by word order or pluralisation.

Some of this can be normalised safely. Some of it carries real meaning.

Take "iphone 17 pro case" and "iphone 17 pro max case". They differ by a single token. A clean-up routine that strips common modifiers or compresses product names too aggressively can make these look far more similar than the actual business reality allows.

Start with canonical text handling

Unicode allows visually identical strings to be built from different underlying code-point sequences. The Unicode Consortium recommends normalising text so canonically equivalent strings compare consistently. NFC is generally the safest choice for ordinary text, because it keeps canonical distinctions intact without applying the heavier compatibility transformations that NFKC and NFKD use.

For keyword pipelines, this is the kind of cleaning that should just happen automatically in the background. Two strings that are visually and canonically identical shouldn't end up as separate observations just because they were encoded differently.

A practical first pass can include:

  • Unicode NFC normalisation;
  • whitespace trimming and collapse;
  • consistent handling of non-breaking spaces;
  • normalisation of obvious formatting artefacts;
  • preservation of the original raw query in a separate column.

That last point really matters. Never throw the source string away. The normalised value is an analytical representation, not a replacement for the original observation.

Lowercasing is usually safe, but not always enough

Plenty of lexical vectorisers lowercase text by default. Scikit-learn's TfidfVectorizer, for instance, lowercases before tokenising unless you tell it not to.

That's often fine for search-query analysis, since "Running Shoes" and "running shoes" normally mean the same thing. But don't confuse case handling with entity handling. Some datasets contain acronyms, product codes or abbreviations where the original surface form is still useful for explanation and QA. A good workflow can cluster on a case-normalised version while still showing the original query to the person reading the report.

Don't strip punctuation blindly

Punctuation can be noise. It can also be part of an entity.

Examples include:

  • C++;
  • .NET;
  • brand names containing punctuation;
  • part numbers;
  • quoted phrases;
  • hyphenated product descriptors.

A generic regex that strips every non-alphanumeric character makes tokenising easier, but it also makes the query less faithful to what the person actually typed. Character n-grams can help here, because they preserve local string structure without needing every token boundary to be perfect. Scikit-learn supports word, character and within-word character n-gram analysers precisely for this kind of choice.

Stop-word removal deserves suspicion

Traditional text pipelines often strip out common words like "the", "for", "to" or "with" to shrink vocabulary size. That's fine for long documents, but search queries are short, and a small word can change the whole task.

Compare:

  • "software for agencies";
  • "software agencies";
  • "how to cancel subscription";
  • "cancel subscription".

In a long document, some function words barely affect topical classification. In a three- or four-word query, removing just one can change the meaning significantly. A sensible default is to leave stop words in unless you've actually tested the effect on your own judgement set.

Stemming and lemmatisation solve a different problem

Lemmatisation maps inflected forms back to a base form using linguistic rules or models. SpaCy, for example, has a lemmatiser component that can use language-specific rules, lookup tables or trainable approaches.

That's helpful when the difference between "run", "running" and "runs" doesn't matter for clustering. But again, justify the transformation by the decision downstream. In ecommerce, singular and plural terms sometimes map to genuinely different page types. In technical search, what looks like a morphological variant might actually be part of a product or command name.

Treat linguistic normalisation as an analytical feature you can add, not an irreversible rewrite of the underlying data.

Exact duplicates and near-duplicates are not the same thing

Easy distinction to lose sight of.

An exact duplicate is the same query once you apply your agreed canonical normalisation. It's generally safe to collapse those repeated rows, as long as you aggregate their metrics appropriately.

A near-duplicate is a different string that may or may not represent the same task. For example:

  • "best crm software" / "best crm softwares";
  • "cheap crm software" / "affordable crm software";
  • "crm software uk" / "crm software";
  • "crm software pricing" / "crm price".

Near-duplicates belong in the clustering or review layer. If preprocessing silently merges them, you lose the chance to check whether the difference actually mattered. We go into this further in keyword deduplication versus clustering.

Preserve commercial modifiers

One of the most common preprocessing mistakes is removing words simply because they occur often across the dataset.

In SEO, those frequent modifiers can be exactly what separates useful page groups:

  • cheap;
  • best;
  • near me;
  • reviews;
  • price;
  • installation;
  • replacement;
  • commercial;
  • UK.

A model might treat these as semantically secondary to the main noun. A practitioner might see them as decisive, because they change page type, funnel stage or geography entirely. Before stripping high-frequency terms, check whether they're acting as stop words or as genuine operational modifiers in your specific corpus.

Keep multiple representations when they answer different questions

You don't have to force one cleaned string to serve every stage of the pipeline.

A robust keyword table can hold onto:

  • raw_query for audit and reporting;
  • canonical_query for deduplication;
  • lexical_query for TF-IDF or n-gram features;
  • embedding_query for semantic encoding;
  • entity fields for brands, products, locations or attributes;
  • source metadata such as volume, country and ranking URL.

That avoids a common design mistake: treating preprocessing as one destructive cleaning function that every downstream component just has to accept.

Use a transformation log

In a production system, every automated transformation should be explainable after the fact.

If "best noise-cancelling headphones UK" becomes "best noise cancelling headphones uk", that's probably harmless. If it becomes "noise headphones", something important just got lost. Store enough to answer:

  • what changed;
  • which rule changed it;
  • whether the original row got collapsed into another;
  • which metrics were aggregated;
  • which representation was actually clustered.

This is especially useful when a practitioner questions an unexpected cluster. There's no point debugging the clustering algorithm if the error was introduced two stages earlier.

A conservative preprocessing sequence

  1. Preserve the raw row. Never overwrite the original keyword.
  2. Normalise text encoding and whitespace. Fix purely technical inconsistencies first.
  3. Create an exact-deduplication key. Collapse only observations that meet a clearly defined equivalence rule.
  4. Apply lexical processing separately. Lowercasing, n-grams or lemmatisation should create features, not destroy the source.
  5. Extract important entities and modifiers. Keep them available as explicit evidence.
  6. Generate semantic representations. Embed the most appropriate text form for the model.
  7. QA the changes. Sample transformed rows from brands, product codes, locations and high-value categories.
Practitioner rule: if a preprocessing step makes the dataset tidier but makes it harder to explain why two real queries are different, it's probably too aggressive.

How to test whether cleaning actually helped

Don't judge a preprocessing rule by how many rows it removes. Evaluate it against a small set of cases where the correct distinction genuinely matters. Watch for two kinds of error:

  • false collapse: meaningfully different queries become indistinguishable;
  • false separation: formatting or encoding differences keep equivalent queries apart.

Then compare how the downstream clusters behave before and after the transformation. If a cleaning rule improves consistency without erasing important modifiers, keep it. If it just produces prettier data and worse decisions, drop it.

ClusterIQ Conclusion

Keyword preprocessing isn't administrative housekeeping. It decides what information the clustering system is even allowed to see.

The safest approach is conservative and reversible: normalise technical inconsistencies, preserve raw values, keep exact duplicates distinct from near-duplicates, and hang onto commercially important modifiers for later stages.

A strong clustering pipeline doesn't start by making every query look the same. It starts by deciding which differences are genuinely meaningless, and which ones a practitioner might need later.

Related ClusterIQ analysis

For specific preprocessing failure modes, see stemming and lemmatisation, character n-grams and acronym handling.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.