Versioning a topical map: how ClusterIQ can evolve without rewriting historical SEO decisions
Topics, models and site structures change. Learn how stable IDs, lineage and run versions let ClusterIQ evolve while preserving the history behind old page and roadmap decisions.

Farky Rafiq
Founder of ClusterIQ

A topical map should be a living document. In a healthy SEO strategy, new products launch, search intent shifts, pages are merged, and the underlying AI models we use for analysis get smarter. However, if every update simply overwrites your previous work, you lose the ability to explain why you made certain decisions six months ago. Historical reports become impossible to reconcile, and your content roadmap loses its foundation.
For ClusterIQ to be truly useful for long-term planning, it needs versioning and lineage. This ensures your topic map can evolve and improve without pretending the past never happened.
Separate topic identity from cluster number
If you have ever exported a few thousand keywords from Ahrefs and run them through a basic script, you will know that "Cluster 27" is not a permanent identity. If you run that same algorithm again tomorrow with slightly different data, the numeric labels will likely change, even if the keywords inside the group are almost identical. To build a durable strategy, we must use stable topic IDs that persist as long as the underlying group remains recognisably the same.
Every analytical run should have a version
To make your SEO structure reproducible, every run needs a clear audit trail. We should store the source data snapshot, the preprocessing version, the specific embedding model used, and the clustering parameters. Recording the software version and a timestamp ensures that when a client asks why a specific group of keywords shifted, you can point to the exact configuration that caused it.
Lineage connects old clusters to new ones
When your data evolves, you need to know how the old structure relates to the new one. Useful lineage events include topics being continued, split, merged, created, or retired. By looking at the overlap of keywords and entities between two versions, we can automatically infer these relationships, keeping your reporting consistent even as the map grows.
Worked example: one topic splits
Imagine an early ClusterIQ run where you have one broad topic for "keyword clustering methods". As you add more data from Search Console or Semrush, the corpus grows. In the next version, that single group might separate into four distinct clusters: density clustering, graph clustering, hierarchical clustering, and evaluation metrics. The original broad topic should not just vanish; it becomes a historical predecessor or a parent linked to these new, more specific children. This allows you to track performance from the broad original intent down to the new granular pages.
Human-approved labels need versioning too
Sometimes an SEO will rename a cluster to make it more descriptive for a content brief without changing the keywords inside it. We need to store both the AI-generated label and the human-approved one, along with who changed it and when. This preserves continuity in your monthly reporting so stakeholders aren't confused by sudden naming changes.
Page mappings have their own history
A cluster might move from Page A to Page B after a content consolidation project. It is vital to remember that the topic itself hasn't changed, only the URL assigned to it. By versioning the mapping relationship separately from the cluster membership, you can track how your site architecture has evolved to better serve those topics.
Business priority should not rewrite topic history
A topic can become a high priority for a business without becoming a different topic. Whether a cluster is "in progress" on your roadmap or "strategically vital" for Q4, these status fields should be attached to the stable topic ID as temporal history. This prevents your project management data from being wiped out during a technical refresh.
Use ARI/NMI for run-level change
When comparing two different versions of a keyword set, metrics like ARI and NMI are incredibly helpful. They allow us to quantify how much the overall structure has changed between versions. Lineage then provides the "why" by explaining the specific splits and merges that make up that number.
Use Jaccard for local lineage
To track how individual clusters evolve, we compare the member sets of old clusters with new candidates. A high Jaccard overlap suggests the topic is continuing as before, while two substantial overlaps might suggest a split. We then use entity and label evidence to confirm this interpretation before committing the change.
Do not auto-accept every lineage match
Algorithms are great at seeing set overlaps, but they can miss a subtle change in page purpose or a shift in search intent. For high-value topics, a practitioner should always have the final say to confirm a split or merge. This human oversight ensures the topical map remains aligned with the actual business goals.
Historical reports should use the version that existed then
If a report from 2026 grouped topics in a specific way, rewriting that report with a 2027 structure would change the meaning of your historical data. ClusterIQ should offer both an "as-reported" historical view and a "restated" view that uses the current lineage. Keeping this distinction explicit prevents confusion during year-on-year reviews.
Versioning enables safe model upgrades
When a better embedding model becomes available, you can create a candidate version of your topic map without touching your live production data. This allows you to compare the two versions, see the improvements, and migrate your strategy deliberately rather than being forced into an immediate change.
Overrides need replay logic
Some human decisions remain valid regardless of how the model changes. For instance, if you have decided that "Series 6" and "Series 8" products must stay in separate groups for legal or branding reasons, that rule should be durable. We replay these human constraints transparently over new model runs, marking them as manual overrides rather than model discoveries.
Topic IDs make roadmaps durable
By referencing a stable topic ID in your tasks and forecasts, your project management links won't break every time you refresh your keyword research. Whether you are managing a few hundred keywords or a few thousand, this stability is what turns a one-off analysis into a reliable system.
Governance should define when a new version is created
Not every tiny edit requires a new version. Major triggers might include a significant data refresh, a change in the embedding model, or a large manual restructuring of the site taxonomy. Establishing these rules helps keep the version history clean and meaningful.
Keep version differences understandable
An experienced SEO should be able to ask what exactly changed between two versions of a map. ClusterIQ answers this by highlighting new or deleted queries, splits, merges, and changes in page mapping. This level of detail makes the evolution of a site's SEO strategy completely transparent.
ClusterIQ principle: a topical map is a versioned model of the site and search landscape. Preserve its history so every future decision remains explainable.
ClusterIQ Conclusion
Versioning transforms a topical map from a temporary export into a durable analytical asset. By using stable IDs and tracking lineage, we allow the model to grow alongside the business. This approach ensures that membership changes, label edits, and page mappings are all recorded, so you never lose the context behind your previous SEO work.
Define approval rules for structural changes
Some changes to a topic map are minor, like refining a label for clarity. Others are significant, potentially altering page owners or reporting structures. ClusterIQ classifies these changes by impact. High-consequence changes, such as a high-value cluster splitting into three, should require practitioner approval. By recording whether a change was driven by new data, a model upgrade, or a manual override, we maintain a clear understanding of the system's evolution.
A practical lineage model for the database
In practice, we treat the topic as a permanent entity. A central table holds the stable ID, while a version table records the specific labels and configurations at a point in time. Relationships connect keywords to these versions, and lineage records link predecessors to successors. This structure makes auditing simple: you can easily see which keywords belonged to a topic on a specific date and which human decisions were active at the time.
Use candidate versions as a safe experimentation layer
Before committing to a new structure, you can generate a candidate version. This acts as a sandbox where you can inspect how a new model or fresh data from Ahrefs would impact your current map. Once you have reviewed the splits, merges, and global changes, you can promote the candidate to production. This ensures that your live reports and roadmaps remain stable while you continue to innovate and improve your analysis.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Keyword cannibalisation: how clustering helps find overlap without inventing a problem

Using keyword clusters to design an SEO taxonomy without letting search volume run the site
