Skip to main content
All articles
Clustering
15 June 2026 4 min read

Normalized Mutual Information for SEO: comparing cluster structures with different labels

Normalized Mutual Information compares how much information two cluster assignments share, regardless of label names. Learn where it helps SEO evaluation and where it can mislead.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Two side-by-side cluster assignments for the same query nodes, with different colour groupings connected to show shared structure despite arbitrary labels.

You have clustered 1,500 keywords from Ahrefs, changed a model or setting, and now the groups look different. Before rebuilding the content plan, you need to know how much actually changed.

Normalised Mutual Information, or NMI, compares two clusterings without requiring their labels to match. It measures how much knowing a query’s group in one run tells you about its group in another. For ClusterIQ evaluation, it is useful when comparing model versions, benchmarks and parameter changes.

What mutual information means here

Mutual information measures dependence between two sets of cluster assignments, also called partitions. Identical memberships with different cluster numbers share a lot of information; unrelated assignments share little.

Normalisation scales the result to make comparisons across different clustering problems easier.

Why label names do not matter

Cluster IDs are arbitrary. “Garden office insulation” might belong to cluster 5 in one run and cluster 27 in another.

NMI compares memberships, not those numbers. That makes it useful for checking repeat-run stability and testing whether model updates introduce unexpected changes.

Worked example: a model upgrade

Suppose ClusterIQ upgrades the sentence embedding model, which represents query meanings numerically, while keeping the keyword list and downstream clustering settings fixed.

An NMI of 0.91 between the old and new assignments suggests substantial shared structure. But “is 0.91 good?” is not the most useful question.

Ask which memberships changed, whether benchmark results improved, and whether the changes lead to better page decisions. A stable overall structure does not automatically mean better content briefs.

NMI and ARI describe change differently

Adjusted Rand Index, or ARI, looks at agreement between pairs of queries. NMI looks at information shared between the assignments.

They often move together, but can disagree when cluster sizes or split-and-merge patterns differ. Using both in ClusterIQ regression testing would make that disagreement a prompt to inspect the changed groups.

NMI is symmetric

Swapping the two clusterings does not change the comparison. Neither needs to be treated as the “correct” answer, which suits comparisons between competing models.

Benchmark comparison needs context

NMI can compare model results with human labels when those labels form a clean partition: each query belongs to one group.

If the benchmark allows overlapping topics or several valid labels per query, forcing it into one partition can lose important ambiguity.

ClusterIQ's benchmark design should therefore determine whether NMI fits the evaluation task.

High NMI can coexist with important errors

Suppose 95% of an ecommerce keyword set stays unchanged, but model-number queries for an important product family are incorrectly merged. The overall structure may still look very similar.

Pair global NMI with local cluster lineage, meaning a record of how individual groups changed, and review high-value queries. Those checks matter before assigning target URLs or updating briefs.

Low NMI can reflect intentional improvement

An old model may produce broad, mixed groups. A new one may separate intent and entities more accurately, lowering NMI because the structure genuinely changed.

The previous partition is a baseline, not truth. Human benchmark results and downstream page decisions determine whether that change is useful.

Cluster count can influence interpretation

Partitions with very different numbers of groups can still share substantial information, particularly when one refines the other.

NMI may remain fairly high when broad parent groups split into coherent children. That can be desirable for a multi-level ClusterIQ topic map, rather than evidence of unwanted instability.

Use contingency tables for explanation

A contingency table counts how members of each old cluster are distributed across the new clusters. It makes changes visible:

  • one old cluster maps mostly to one new cluster;
  • one old cluster splits into three children;
  • three old clusters merge into one broader group.

These lineage events give you a more useful review list for a content plan than the NMI number alone.

Noise needs explicit handling

As with ARI, density-based clustering can complicate interpretation when thousands of unrelated outliers share one “noise” label.

For ClusterIQ evaluation, comparisons could include and exclude noise, depending on the question being tested. Make that choice explicit in the reporting.

Use NMI as a regression indicator

A useful approach is to establish the normal NMI range for small, expected parameter changes on a frozen benchmark keyword set.

If a code or model update suddenly produces a much lower value, it can fail quality assurance until someone reviews the membership differences. This focuses review time on unexpected changes rather than checking every group afresh.

Do not use NMI as a public “quality score”

NMI measures similarity between two partitions. Without a meaningful reference partition, it says nothing about whether either clustering is good for SEO.

Use it in model evaluation and stability reporting, not as a headline score without context. A client needs to understand the implications for their site, not just see a reassuring number.

Practitioner principle: NMI tells you how much structural information two clusterings share. It does not tell you which one better serves the site.

ClusterIQ Conclusion

NMI is a useful structural comparison metric for ClusterIQ evaluation. Alongside ARI, local lineage, human benchmarks and page-level review, it makes clustering changes measurable and helps focus practical checks. It should support decisions about content and URLs, not turn mathematical similarity into an SEO verdict.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.