Skip to main content
All articles
Clustering
24 September 2026 5 min read

Adding new keywords to existing clusters without rebuilding everything

New keywords do not always require a full recluster. Learn when fixed assignment is safe, when new topics should stay unassigned and when the structure needs rebuilding.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Diagram showing new keyword points assigned to existing clusters, held for review, or forming a separate potential new topic.

You have grouped your keywords, mapped the topics to pages and built a report. Then a fresh Search Console export arrives. Do you need to start again, or can you add the new queries to the groups you already have?

It is a recurring problem. New products launch, markets shift and research tools turn up more keywords. Rebuilding the whole clustering process with every update can be costly and makes results harder to compare over time.

But dropping every new keyword into an existing group has its own risk: those groups may no longer reflect what people are searching for.

The useful distinction is between fitting new keywords into an existing structure and working out whether that structure needs to change.

Assignment and reclustering are not the same operation

Suppose you grouped 2,000 keywords last month and receive 150 new queries today.

You have two broad options:

  • Assignment: keep the existing clusters fixed and decide where the new queries fit.
  • Reclustering: combine the old and new keywords, then let the grouping process change the cluster structure.

Assignment costs less to run and keeps your reporting consistent. Reclustering takes more work, but it can uncover new topics and reorganise existing ones.

When assignment is appropriate

Adding new keywords to existing clusters works well when:

  • the new batch is small compared with the original dataset;
  • the market has not changed dramatically;
  • most new queries are likely to be variations on topics you already cover;
  • you need stable cluster IDs for reporting;
  • you have a clear way to reject weak matches.

That last point matters. “None of the above” must be a valid result. If a keyword does not convincingly fit an existing cluster, the system should leave it unassigned.

Nearest-neighbour assignment is the simplest approach

If your existing keywords have embeddings, numerical vectors that represent their meaning, you can create an embedding for each new keyword using the same model. You then compare it with the existing vector index, the searchable collection of those representations.

For each new keyword, you can check:

  • its nearest individual neighbours, meaning the existing keywords with the most similar vectors;
  • which clusters those neighbours belong to;
  • how closely it matches representative keywords from each cluster;
  • whether its closest neighbours agree on a cluster.

If the closest matches consistently belong to one cluster and pass a similarity threshold calibrated for your data, assignment can be straightforward.

If those matches are spread across several clusters, send the query for review rather than automatically choosing whichever cluster has the most neighbours.

Cluster exemplars can reduce the search space

You do not always need to compare a new query with every keyword you have collected.

Each cluster can keep a set of representative vectors, such as:

  • a centroid, the average vector, or a medoid, an actual member chosen to represent the group;
  • several high-confidence exemplars, meaning keywords that clearly belong in the cluster;
  • representative combinations of entities, such as brands, products or locations;
  • boundary exemplars, which represent areas where clusters commonly overlap.

First, compare the new query with these cluster-level representatives. Then compare it with individual keywords only within the most promising clusters.

This two-stage search is faster and makes the assignment process easier to explain.

HDBSCAN makes the assignment/rebuild distinction explicit

HDBSCAN, a clustering algorithm, provides an approximate_predict() function for assigning new, unseen data points to clusters in an already fitted model.

Its documentation highlights an important limit: this is not the same as running HDBSCAN again with the new points included. The prediction function assigns points to clusters that already exist. It does not discover new clusters in the incoming data.

For an SEO workflow, that distinction is essential. Finding a home for a new keyword is not the same as checking whether you need a new topic group.

Incremental learning can change the model

Some algorithms can update the model itself as new data arrives, rather than just assign keywords to fixed groups. This is called incremental learning.

For example, scikit-learn's MiniBatchKMeans offers partial_fit(), which updates its estimates of cluster centres using small new batches of data.

That is useful if your intended approach is to divide keywords around centres that are continuously updated.

The trade-off is that those centres can move. As a result, historical assignments may no longer mean quite what they meant before the update.

If continuity matters, save each model state as a separate version rather than quietly changing one permanent clustering model.

New queries can reveal a missing cluster

The main weakness of fixed assignment is that it interprets everything through yesterday's topic structure.

Imagine an ecommerce dataset of 2,000 keywords, created before the business introduced a new product category. A month later, you collect a few hundred queries about that category.

Nearest-neighbour assignment might spread them across several existing clusters because there is no suitable group for them yet.

Signals that should trigger reclustering include:

  • a growing number of unassigned queries;
  • new queries consistently forming a dense neighbourhood, meaning they are closely related to one another;
  • high-confidence assignments causing a cluster to expand unusually;
  • new entities, such as products or brands, that were absent from the original dataset;
  • large shifts in search demand or the way the product range is organised.

Use a holding area for genuinely new structure

A practical approach is to give incoming keywords one of three statuses:

  • assigned to an existing cluster;
  • review required;
  • unassigned / new-topic candidate.

You can periodically cluster the new-topic candidates among themselves. If a stable group emerges, the next full rebuild can include it explicitly.

This gives new demand somewhere to go without forcing it into categories that no longer fit.

Stable cluster IDs need an identity strategy

A full reclustering can change both the membership of groups and the numbers used to identify them.

If your reports or content management system mappings depend on cluster IDs, you need a way to connect the old structure with the new one.

Useful evidence includes:

  • membership overlap, or how many keywords the groups share;
  • representative queries;
  • overlap in entities;
  • similarity between cluster labels;
  • similarity between cluster centroids.

Using that evidence, you can match a new cluster to its most likely historical counterpart, record a split into descendant clusters, or mark a group as genuinely new.

This is about lineage, tracking how groups develop over time, rather than simply giving them consistent names.

Do not update embeddings with a different model

For similarity scores to remain comparable, new queries must use the same representation as the keywords already in your index.

If you change the embedding model, rebuild the existing vectors too, or create a separate model version.

Mixing vectors from different models in one similarity space is not a valid incremental update. You cannot assume those models represent meaning in the same way.

This is why the versioning practices covered in reproducible keyword clustering matter in day-to-day operation.

A practical refresh policy

You do not have to choose between assigning every update and rebuilding every time. You can combine low-cost assignment with periodic reviews of the structure.

For example:

  1. Create embeddings for new keywords using the current model.
  2. Search for matches in the existing vector index.
  3. Automatically assign high-confidence matches.
  4. Send ambiguous cases for review.
  5. Keep weakly matched terms in the holding area as new-topic candidates.
  6. Monitor how many keywords are in that holding area and how they relate to one another.
  7. Trigger a full reclustering when a defined threshold is reached.
  8. Map the new clusters to their historical lineage.

Your trigger could be based on elapsed time, the volume of new keywords or signs of structural change. There is no single refresh schedule that suits every dataset.

Operational principle: assignment preserves yesterday's model. Reclustering asks whether yesterday's model is still the right one.

What to monitor between rebuilds

Useful indicators include:

  • the percentage of new queries assigned automatically;
  • the percentage sent for review;
  • the percentage left unassigned;
  • average assignment confidence;
  • the rate at which new entities appear;
  • clusters receiving unusually high proportions of new terms;
  • the size of the unassigned pool and how closely its keywords relate to one another.

If more queries need review or remain unassigned, that can be an early warning that your topic structure is becoming outdated.

ClusterIQ Conclusion

Adding keywords to an existing clustering system is not just a smaller repeat of the original job. You need to decide whether you are filling out familiar topics or making room for something new.

Fixed assignment is valuable because it is inexpensive, stable and well suited to familiar variations. It should always allow weak matches to be rejected, so genuinely new topics are not squeezed into old groups.

Periodic reclustering is still necessary because markets and datasets change. A useful system supports both: ongoing assignment for continuity and versioned rebuilds to discover topics and rethink the structure.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.