Inter-annotator agreement for SEO: measuring whether humans agree before judging the clustering model
A model benchmark is only as useful as the human labels behind it. Learn how inter-annotator agreement exposes ambiguous SEO cases before model accuracy is interpreted.

Farky Rafiq
Founder of ClusterIQ

It is incredibly tempting to dismiss an automated model as "wrong" the moment a single reviewer disagrees with its output. However, the more revealing question for any SEO is whether two experienced humans would actually agree with each other when faced with the same set of keywords.
Inter-annotator agreement is a method used to measure the consistency of human judgements. For ClusterIQ, this process is vital. it helps us distinguish between genuine model errors and those murky, ambiguous SEO cases where there is no single "correct" answer. It also highlights when our own benchmarking instructions are not clear enough for the team to follow.
Why human labels are not automatically ground truth
If you give two senior SEOs a list of a thousand keywords from Ahrefs or Semrush, they might reach different conclusions on how to group them. This usually happens because they are solving slightly different problems in their heads.
One specialist might be looking for topical relevance, while another is thinking about whether a single URL can realistically rank for both terms. A third might have specific knowledge of a client's business rules that overrides standard SEO logic. Before we can judge a model, we have to make the annotation task completely explicit.
Define the task precisely
Vague instructions lead to messy data. We need to be specific about what we are asking the reviewers to decide:
- Are these two queries semantically related?
- Should these keywords belong to the same SEO cluster?
- Should one page target both, or do they require separate URLs?
- Do these queries share the same search intent and page type?
- Do these terms refer to the exact same entity?
Each of these questions can result in different, yet valid, labels for the same pair of keywords.
Use independent first-pass annotation
If your team sits in a room and discusses every keyword before labelling it, the final agreement rate will look perfect but tell you very little. A more robust workflow involves having reviewers label a sample of a few hundred keywords independently. We then measure the level of disagreement before moving to an adjudication phase. ClusterIQ is designed to preserve these original independent labels alongside the final agreed version.
Cohen's kappa is useful for categorical judgements
When dealing with "either/or" decisions, like whether to use the same page or separate pages, we often use Cohen's kappa. This statistical measure compares the agreement between two people while adjusting for the agreement that might happen just by chance.
While useful, it is not a universal grade of quality. The balance of your categories can skew the result, so it should be viewed as one part of a broader diagnostic toolset.
Worked example: same page or separate page?
Imagine two SEOs reviewing 200 query pairs. They agree on 170 and disagree on 30, giving a raw agreement rate of 85%.
The 30 disagreements are actually more valuable than the 170 agreements. By digging into those conflicts, we often find patterns. Perhaps the disagreements are clustered around:
- Keywords with the same topic but subtly different intent;
- Complex product variants;
- Local modifiers versus national intent;
- Pricing queries versus comparison guides;
- Short, ambiguous "head" terms.
These patterns show us exactly where the annotation guide, and the ClusterIQ model itself, need more refinement.
Disagreement can improve the rubric
If your team is consistently interpreting "same cluster" in different ways, it is time to rewrite the guidelines. A good rubric should include the specific business decision being modelled, clear examples of "yes" and "no" cases, and instructions on how to handle business-specific exceptions or ambiguous terms.
An uncertain label can be valuable
Forcing a "yes" or "no" choice can create a false sense of certainty. ClusterIQ allows for an "uncertain" state. If both the humans and the model are struggling with the same boundary cases, it suggests these keywords should be routed for manual practitioner review rather than being automated blindly.
Measure agreement by failure family
A single percentage score can hide a lot of detail. It is much more useful to report agreement across different categories, such as:
- Simple paraphrases;
- Shifts in intent;
- Entity variants;
- Product attributes and specifications;
- Geographic locations;
- Brand vs non-brand modifiers.
This approach aligns with the principles found in ClusterIQ's human benchmark workflow.
Adjudication should preserve the reason
When reviewers discuss a disagreement and reach a final decision, we must record why the label changed. These reasons are gold dust. they can be turned into new training examples, new model features, or specific workspace rules for a client's content plan.
Do not train blindly on adjudicated labels
A human decision is often specific to a particular site architecture. ClusterIQ distinguishes between general semantic rules and specific client decisions. This prevents a benchmark for one project from accidentally becoming a universal policy that ruins the reporting for another.
Compare model performance with human agreement
If two expert SEOs only agree 82% of the time, it is unrealistic to expect a model to hit 99% agreement with either of them. The human agreement rate provides the necessary context for interpreting model scores.
Use disagreement to design the interface
Frequent disagreement is a signal that the user needs more data to make a call. If experts need to see SERP evidence or entity data before deciding, ClusterIQ should surface that information directly within the review interface.
Inter-annotator agreement also tests documentation
If agreement rates drop when a new junior SEO joins the project, it might not be a reflection of their skill. It often indicates that your task documentation is unclear. Treat your rubric and example sets with the same version control you would apply to the model itself.
Keep the benchmark alive
The benchmark should not be a frozen set of easy examples. Production overrides and difficult manual reviews should be adjudicated and fed back into the system. This ensures ClusterIQ stays aligned with the real-world decisions you face every day.
Practitioner principle: before asking whether the model agrees with humans, establish whether humans agree with one another and whether they are answering the same question.
ClusterIQ Conclusion
Inter-annotator agreement brings rigour to human evaluation. By using independent annotation and detailed disagreement analysis, ClusterIQ builds better benchmarks and more intuitive review interfaces. The aim is not to pretend ambiguity doesn't exist, but to identify exactly where it lies so we can handle it intelligently.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

How to evaluate keyword cluster quality: metrics, stability and human review

Keyword clustering vs topic clustering: what’s the difference?
