Jaccard similarity for SEO: measuring SERP and cluster overlap without oversimplifying
Jaccard similarity is a simple way to compare sets such as SERP URLs or cluster members. Learn where it helps SEO analysis and where the simplicity hides important differences.

Farky Rafiq
Founder of ClusterIQ

If you have ever looked at two different search results pages and wondered exactly how much they have in common, you are essentially asking for a similarity measurement. In SEO, we often need a clear way to quantify how much two groups of things overlap, whether those are lists of URLs, groups of keywords, or sets of pages. Jaccard similarity is a straightforward tool that answers this specific question: how much do these two sets share?
It is a concept simple enough to explain to a client over a quick call, yet robust enough to handle thousands of rows of data from Ahrefs or Search Console. However, its simplicity is a double-edged sword. While it gives us a clean number, it treats every item in your list as having equal value, ignoring things like search rank or the commercial intent of a specific keyword.
What Jaccard similarity measures
In technical terms, Jaccard similarity is calculated by taking the size of the intersection of two sets and dividing it by the size of their union. To put that into a practical SEO context: if Query A and Query B both return ten URLs, and six of those URLs appear on both pages, the intersection is six. The union (the total number of unique URLs across both sets) is fourteen. Dividing six by fourteen gives us a Jaccard similarity of roughly 0.43.
The resulting score sits between zero and one. A zero means there is no overlap at all, while a one means the sets are identical. Anything in between represents a partial match.
Why it is useful for SERP overlap
Measuring SERP overlap is perhaps the most common use for this metric. When two different search terms consistently return the same URLs, it is a strong signal that search engines view those queries as serving the same intent. This data helps you decide whether to target two keywords with a single page or create separate content for each.
Jaccard is often better than just counting shared URLs because it accounts for differences in set sizes. If one search term returns ten results and another returns twenty, the union-based calculation provides a more balanced perspective. This makes it a perfect partner for ClusterIQ's SERP overlap clustering guide.
It can compare cluster membership too
When you are managing large keyword sets, you might run a clustering process, tweak your settings, and run it again. Jaccard similarity allows you to compare how stable your clusters are between these versions. If a cluster had 100 keywords in the first run and 110 in the second, with 90 keywords appearing in both, Jaccard gives you a quick way to see how much the group has actually changed.
This is far more effective than looking at cluster IDs, which often change randomly between runs. In ClusterIQ, using set-based comparisons helps you see if a topic is structurally consistent even if the software has assigned it a new label.
Jaccard ignores rank
One major caveat for SEOs is that Jaccard does not care about position. If two queries share the same six URLs, but those URLs are in positions 1 through 6 for the first term and positions 5 through 10 for the second, Jaccard treats them as identical. In reality, these two SERPs behave very differently for a user.
If the order of results is vital for your task, you should supplement Jaccard with other data, such as top-three overlap or rank-weighted scores. Always keep the raw list of shared URLs handy so you can see the reality behind the summary percentage.
It ignores the importance of individual items
When looking at keyword clusters, a basic Jaccard score treats every keyword as equal. A high-intent, high-volume "money" term carries the same weight as a tiny long-tail variation. While this is fine for checking the structural health of a cluster, it does not tell you the business impact of a change.
You can extend the formula to weight items by volume or value if needed, but keeping the standard unweighted version provides a clean, easy-to-understand baseline for your reporting.
Worked example: two page candidates
Suppose you are using ClusterIQ to evaluate two different landing pages for a specific keyword cluster. Page A ranks for 40 keywords in that cluster, while Page B ranks for 25. They share 18 keywords in common. A Jaccard score can quickly tell you the extent of this overlap.
However, the metric alone won't make the final decision for you. You still need to look at which keywords are shared, the intent behind them, and whether one page is clearly outperforming the other. The score highlights a relationship that needs your attention; it doesn't automatically mean you should merge the pages.
Jaccard is strong when the set definition is strong
The number you get is only as good as the data you put in. When comparing SERPs, you must ensure your data collection is consistent regarding location, device type, and language. If you compare a mobile SERP from London with a desktop SERP from Manchester, the Jaccard score will be skewed by those variables rather than the queries themselves.
Consistency is key. If your rules for what goes into a "set" change, the similarity score becomes impossible to interpret accurately.
Do not use one threshold universally
There is no "magic number" for Jaccard similarity. A score of 0.5 might indicate a very strong relationship when comparing ten search results, but it might mean something entirely different when comparing two massive content clusters. Much like setting semantic similarity thresholds, you need to calibrate your expectations based on the specific dataset you are working with.
Where ClusterIQ can use Jaccard well
Within ClusterIQ, Jaccard serves as a transparent diagnostic tool rather than a hidden "black box" algorithm. It is incredibly useful for checking cluster lineage, identifying potential duplicate content, and analysing how queries map to specific pages. Because the logic is so clear, you can always dig into the data to see exactly why a specific score was generated.
Practitioner principle: Jaccard is excellent at telling you the scale of an overlap. It cannot tell you if that overlap is profitable, intentional, or enough to justify a change in your content strategy.
ClusterIQ Conclusion
Jaccard similarity is a vital, simple metric for any SEO who works with clusters. It provides a clear way to compare SERPs, keyword groups, and page relationships without getting bogged down in overly complex math. Just remember to use it as a starting point alongside other factors like rank, intent, and business value to get the full picture.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Soft clustering and confidence scores: handling ambiguous keywords honestly

Adding new keywords to existing clusters without rebuilding everything
