Skip to main content
All articles
Clustering
24 September 2026 5 min read

Reproducible keyword clustering: how to make every run explainable

Keyword clusters change when data, models or parameters change. Learn how run manifests, versioning, seeds and audit trails make clustering explainable and comparable.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing two keyword-clustering runs linked through a central manifest, with stable clusters preserved and a small cluster split and boundary node change made visible between versions.

If nobody can reproduce a keyword-clustering result, it's basically impossible to trust it, compare it, or improve it.

Two runs on the same task can produce different outputs for all sorts of reasons: the data changed, the embedding model changed, a random seed changed, a UMAP setting changed, or a community-detection parameter moved slightly. Without a versioned record of those choices, every one of those causes just looks the same from the outside: "the clustering changed".

Reproducibility isn't an academic nice-to-have here. It's basic operational hygiene for any clustering system that's actually influencing content, taxonomy or URL decisions.

What actually needs to be reproducible?

A useful clustering run is more than just a CSV of final labels. You should be able to reconstruct:

  • the source dataset;
  • the preprocessing rules;
  • the representation;
  • the similarity or neighbour-retrieval method;
  • the clustering or community algorithm;
  • all the important parameters;
  • random seeds where applicable;
  • manual overrides;
  • the final labels and quality diagnostics.

Miss any of these, and reproducing the exact output later becomes a lot harder than it should be.

Version the data, not just the code

The same code run against a different keyword export can legitimately produce different clusters. Store a dataset fingerprint or immutable source identifier alongside every run. Useful fields include:

  • source filename or job ID;
  • row count;
  • upload timestamp;
  • market and language;
  • hash of the canonical input;
  • preprocessing version.

That lets you tell a model change apart from a data change, which sounds obvious until you're trying to debug a change with neither recorded.

Embedding models need explicit versions

"We used Sentence Transformers" isn't actually reproducible on its own. Record:

  • model name;
  • model revision or checksum where available;
  • vector dimensionality;
  • normalisation setting;
  • exact text representation embedded.

If the model gets upgraded later, keep the old run intact. A new model is a new analytical configuration, not a silent patch to the old one.

Randomness can change the output

Several methods used in keyword clustering have stochastic elements built in. UMAP, for instance, documents its own reproducibility considerations around random state. Community-detection methods can depend on seed or node ordering. K-means commonly uses random initialisation unless you configure it otherwise.

A fixed seed makes comparisons easier, but it doesn't make the method fully deterministic across every environment or library version. The point of recording seeds is to control one known source of variation, not to claim perfect reproducibility everywhere.

Save library versions too

Numerical and machine-learning libraries keep evolving. Changes to defaults, optimisation code or dependencies can shift behaviour even when your own application code hasn't changed at all.

For important runs, record versions of:

  • Python;
  • Sentence Transformers;
  • scikit-learn;
  • UMAP;
  • HDBSCAN;
  • NetworkX;
  • vector-search libraries.

A container or lockfile makes all of this far easier to reproduce later on.

Parameters are part of the result

A cluster output without its parameters is incomplete. For a graph workflow, store:

  • similarity metric;
  • threshold;
  • top-k neighbour count;
  • mutual-neighbour rule;
  • edge-weight formula;
  • community algorithm;
  • resolution;
  • seed.

For HDBSCAN, store min_cluster_size, min_samples, metric and any dimensionality-reduction configuration. The configuration isn't just metadata sitting around the result. It's part of what the result actually means.

Don't overwrite human overrides

Practitioner edits create a second layer of truth. If someone renames a cluster, moves a keyword or merges two groups, keep hold of:

  • original model output;
  • human override;
  • reason;
  • reviewer;
  • timestamp.

A later model run can then be compared against both the old model and the approved human state, rather than just overwriting one with the other.

Reproducibility makes experiments cheaper

Once runs are versioned, testing parameters becomes genuinely useful rather than guesswork. You can compare:

  • model A versus model B;
  • threshold 0.72 versus 0.76;
  • Louvain versus Leiden;
  • with and without entity constraints;
  • original versus revised preprocessing.

Without stable run IDs and configurations, all these experiments just turn into a folder of screenshots and half-remembered impressions.

Store quality metrics per run

Every run should carry an evaluation summary. Possible fields include:

  • number of clusters;
  • cluster-size distribution;
  • noise rate;
  • largest connected component;
  • silhouette or structural metrics where appropriate;
  • judgement-set accuracy;
  • review rate;
  • number of manual overrides.

These don't prove one run is better than another, but they make changes visible instead of invisible.

Use a run manifest

A simple manifest can make the whole system auditable. For example:

{
  "run_id": "2026-09-24-001",
  "dataset_hash": "...",
  "preprocessing_version": "1.2",
  "embedding_model": "...",
  "similarity": "cosine",
  "top_k": 20,
  "threshold": 0.74,
  "community_algorithm": "leiden",
  "resolution": 1.0,
  "seed": 42
}

The exact schema will change over time. The principle shouldn't: the result needs to carry enough information to explain how it was produced.

Reproducible doesn't mean frozen forever

A reproducible system should still keep improving. The difference is that improvements create new versions rather than silently changing the meaning of old results.

That's especially important for a SaaS product. A customer might come back to an analysis weeks later. If the cluster structure has shifted because the default model got updated behind the scenes, the interface should be able to explain that, rather than just quietly showing a different answer.

Stability testing becomes more meaningful

Our guide to evaluating cluster quality recommends testing nearby parameter settings and seeds. Versioned runs make that genuinely practical: you can measure which core groups persist and which boundary queries drift. That turns "the model feels unstable" into something you can actually measure.

Compare runs with a structured diff

Reproducibility gets even more valuable when the system can explain what changed between two runs. A run comparison can report:

  • clusters that persisted with similar membership;
  • clusters that split or merged;
  • keywords that moved between groups;
  • new or removed observations;
  • changes in noise or review rates;
  • configuration differences that could explain the movement.

That makes a model upgrade much easier to evaluate. Instead of just asking whether the new result "looks better", the team can inspect exactly which decisions changed and whether those changes are actually desirable.

Keep historical outputs immutable

Once a run has informed a report or a page-mapping decision, avoid rewriting it in place. Save improvements as a new run and link the two through lineage instead.

That gives practitioners a reliable audit trail and stops a dashboard from silently rewriting the past whenever defaults get updated. Historical reproducibility matters a lot when manual overrides, client approvals or implementation work have already been based on the earlier result.

Engineering principle: if a clustering result affects a business decision, treat its configuration like code: version it, record it and make changes explicit.

A minimum reproducibility checklist

  1. Immutable source data or dataset hash.
  2. Preprocessing version and transformation rules.
  3. Embedding model and revision.
  4. Similarity and neighbour configuration.
  5. Clustering algorithm and parameters.
  6. Random seed where applicable.
  7. Library and runtime versions.
  8. Quality metrics.
  9. Human overrides and labels.
  10. Timestamped run ID.

ClusterIQ Conclusion

Keyword clustering is really a chain of modelling choices. If those choices aren't recorded, the final cluster labels are hard to audit and almost impossible to compare properly over time.

Reproducibility doesn't require every run to stay identical forever. It requires every change to come with an explanation.

That's what turns clustering from a one-off black-box output into a system that practitioners can actually test, challenge and improve.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.