Skip to main content
All articles
Clustering
16 June 2026 4 min read

Adjusted Rand Index for SEO: comparing cluster assignments without being fooled by chance

Adjusted Rand Index compares two clusterings while correcting for agreement expected by chance. Learn how it can test stability, model changes and human benchmarks in SEO.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Two side-by-side clusterings of the same observations, with most group memberships matching and one subgroup changing between partitions.

You rerun a keyword clustering job and the groups look similar. Before updating a content plan or changing page assignments, it helps to know how much actually moved, rather than checking hundreds of keywords by hand.

Adjusted Rand Index, or ARI, measures agreement between two clusterings of the same items. It compares which keyword pairs belong together or apart, then adjusts for agreement expected by chance. For ClusterIQ, it is useful when testing model changes, parameter stability and agreement with benchmarks.

What ARI compares

Imagine clustering the same 1,000 keywords from Ahrefs, Semrush or Search Console twice. For every pair of keywords, each run effectively says:

  • these belong in the same cluster; or
  • these belong in different clusters.

ARI summarises how consistently the runs make those decisions, while correcting for chance agreement. Each complete set of cluster assignments is called a partition.

Why the adjustment matters

The unadjusted Rand Index can look reasonably high partly because most keyword pairs belong to different clusters in both runs. Agreement about keeping pairs apart can dominate the result.

ARI accounts for agreement that could occur under random labellings with comparable cluster structures. A score near 1 indicates strong agreement. Values near 0 roughly indicate chance-level agreement under this adjustment. Negative values mean agreement is worse than expected by chance.

Cluster labels do not need to match

Cluster numbers are arbitrary. Run A might call a topic Cluster 7 while Run B calls it Cluster 31. ARI compares membership relationships, not whether those numbers match.

This matters for ClusterIQ stability analysis, where IDs can change between runs without the underlying keyword groups changing.

Worked example: threshold regression

Suppose ClusterIQ changes a semantic edge threshold from 0.78 to 0.80. This changes the similarity requirement for linking keywords before rerunning community detection on the same 2,000 queries.

An ARI of 0.96 suggests the overall partition is highly stable. An ARI of 0.42 signals much greater structural change, giving you a reason to inspect the groups before refreshing content briefs.

Neither score tells you whether the change is good. It tells you how much membership changed.

Use ARI to compare model versions

Changing the embedding model, which represents keyword meaning numerically, can alter thousands of neighbour relationships. ARI provides a whole-dataset regression measure for comparisons such as:

  • old model clustering versus new model clustering;
  • old preprocessing versus new preprocessing;
  • old graph rule versus new graph rule.

ClusterIQ can then drill into which clusters split, merged or changed membership. That makes the overall score a starting point for review, not the final deliverable.

High ARI can still hide important local changes

If 95% of the dataset stays identical while one revenue-critical cluster changes dramatically, the overall ARI can still look strong. A stable report does not necessarily mean your most important landing-page plan is unchanged.

Pair the score with checks of:

  • changes within high-value clusters;
  • important queries moving between groups;
  • page mapping changes;
  • changes to outliers.

One overall quality score should never replace local inspection.

Low ARI is not automatically bad

A deliberate model improvement may fix many weak assignments, producing a lower ARI against the previous version. The old clustering is not ground truth merely because it came first.

Use a human benchmark and practitioner review to judge whether the new structure supports more useful content plans and page decisions.

ARI can compare a model with a labelled benchmark

When a human benchmark gives each keyword one cluster label, ARI can compare the model's partition with that reference.

This suits a genuine partitioning task: every keyword belongs to exactly one group. It is less appropriate when human labels overlap or allow multiple labels per keyword, because ARI assumes a single cluster assignment for each observation.

Human agreement can also be measured with ARI

Ask two expert reviewers to cluster the same benchmark independently. ARI between their partitions helps reveal whether the task has a clear human answer.

If experts strongly disagree, demanding near-perfect model agreement with one reviewer is not sensible.

Compare like with like

Both lists of labels must cover the same observations, with corresponding keywords aligned. If one run adds 500 keywords, compare the shared keywords first, or use a lineage-aware analysis that tracks groups across the changed dataset.

This keeps model drift separate from data drift: changes caused by the model versus changes in the keyword input.

Noise labels need deliberate treatment

Density-based methods may give many observations a shared noise label, such as -1. ARI then sees those observations as belonging together, although they are not necessarily one coherent cluster.

ClusterIQ can calculate ARI with and without noise, or treat noise cases separately, depending on the question.

ARI complements set-based lineage measures

ARI compares whole partitions. Jaccard similarity compares one cluster's members across versions. Together, they provide global partition stability and local cluster lineage: the overall picture alongside what happened to individual groups.

Use thresholds for review, not truth

No universal ARI value means an SEO clustering is “good”. ClusterIQ can establish empirical ranges from known stable configurations and use large drops as regression warnings, rather than absolute quality grades.

Practitioner principle: ARI measures agreement after accounting for chance. It does not tell you which partition is better for SEO.

ClusterIQ Conclusion

ARI is a strong regression and stability metric for keyword clustering. ClusterIQ can use it to quantify structural change across parameters, model versions and benchmark labels, while retaining local comparisons for the content and URL decisions that matter most.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.