Keyword Clustering and Graph Theory Explained
Learn how graph theory, semantic embeddings and community detection can improve keyword clustering and reveal relationships across large keyword datasets.

Farky Rafiq
Founder of ClusterIQ

You have a spreadsheet of 1,500 keywords and need to turn it into a content plan. Which searches belong on the same page? Which need their own pages? And how should those pages fit together across your website?
Keyword clustering helps answer these questions by grouping searches according to shared meaning, intent or topic.
Traditionally, clustering has often relied on fairly straightforward rules. Keywords might be grouped because they contain similar words, return the same search results, or meet a set level of similarity in meaning.
These approaches are useful, but they can oversimplify how search topics connect. Graph theory offers a way to keep more of those relationships in view.
A graph is a model of connections between individual things. For keyword research, each keyword becomes a node, and a relationship between two keywords becomes an edge.
Rather than treating your keyword list as hundreds or thousands of separate spreadsheet rows, you can start to see it as a network.
Keywords as a network
Imagine you are planning content for a bathroom retailer. Your keyword list includes:
walk-in shower
walk-in shower enclosure
shower enclosure
frameless shower enclosure
frameless shower screen
shower screen
A traditional clustering system might compare pairs of keywords and measure how similar they are. A graph-based approach uses those comparisons to build a picture of the connections across the whole list.
Each keyword becomes a node. If two keywords are sufficiently related, an edge connects them. That edge can also have a weight, a value representing how closely related the keywords are.
Across a dataset of 1,500 keywords, thousands of these connections can form a network.
Some keywords will have many strong connections. Others will sit at the edges of a topic, with fewer or weaker links. Some groups will be closely connected to one another, forming communities within the network.
Those communities can provide a natural basis for keyword clusters.
Semantic embeddings make this possible
To build a useful keyword network, you need a way to recognise related meanings, not just matching words. This is where semantic embeddings have become particularly useful.
Embedding models convert language into numerical representations called vectors. Keywords with similar meanings tend to sit near one another in the numerical space created by the model.
That gives us a way to calculate similarity in meaning mathematically.
A common measure is cosine similarity, which checks how closely aligned two vectors are. A high cosine similarity score suggests that two keywords are related in meaning.
For example, an embedding model should recognise the close relationship between:
“cheap running trainers”
“affordable running shoes”
The phrases share relatively few words, but describe much the same thing. Recognising that relationship is a major improvement over relying on word matching alone.
Once you have calculated these similarity scores, you can use them to build a graph. The keywords become nodes, and the scores determine which nodes connect to one another.
From similarity to communities
Building the network is only part of the job. You still need to identify meaningful groups within it.
Community detection algorithms help with this step. Algorithms such as Louvain look for groups of nodes that connect more strongly to one another than to the rest of the network.
For SEO, that means looking for groups of keywords that appear to describe a coherent topic.
You do not have to tell the algorithm beforehand that you want, say, 50 clusters. Instead, the structure of the graph helps determine where natural communities exist.
This can be especially useful when you have a large keyword list and the boundaries between topics are not obvious.
For example, a broad topic such as “bathroom furniture” may contain distinct communities around:
vanity units
bathroom cabinets
mirrored cabinets
wall-hung storage
freestanding bathroom furniture
Smaller subgroups may also emerge within those communities.
The resulting structure can potentially reflect the search landscape more closely than grouping keywords only by the words they share.
Why graph theory matters for SEO
The value here is not the mathematics for its own sake. It is the extra structure you can uncover and use in your planning.
SEO topics rarely fall into neatly separated boxes. A keyword can belong strongly to one topic while still having meaningful connections to another. Some keywords act as bridges between different parts of a subject.
Graph models are well suited to showing these overlapping relationships.
Take “shower bath”. It could sit between a broad bath cluster and a shower-related cluster. Assigning it to just one category may hide useful information about its relationship with the other.
In a graph, its connections to both topics can remain visible, even if you ultimately assign it to one cluster.
Looking at these connections can help you identify:
Primary topics. Closely connected groups can point to major areas of search demand.
Subtopics. Smaller communities within a larger cluster can suggest useful page sections or supporting content.
Bridge terms. Keywords that connect separate communities can show where topics overlap.
Outliers. Keywords with weak connections may have a different intent, be irrelevant to your project, or represent opportunities worth investigating.
Central concepts. Nodes with unusually high numbers of connections can indicate important concepts within a topic.
These patterns can inform more than your keyword groups. They can help you plan your site's information architecture, internal links, category structure and content.
Graphs should not replace search intent
There is an important limit to all of this: two keywords can be related in meaning without belonging on the same page.
Consider these searches:
“best running shoes”
“running shoe shop”
They are clearly related, but the people searching may want different things. One may be comparing options, while the other may be looking for somewhere to buy.
Google's search results also offer valuable behavioural evidence about whether queries are treated similarly. That is why a graph should be part of a broader clustering framework, not the final answer on its own.
Different signals contribute different information:
Semantic embeddings help identify meaning.
SERP overlap, the extent to which searches return the same ranking pages, provides evidence of ranking similarity.
Intent classification helps distinguish informational, commercial and transactional searches.
Graph algorithms can organise those signals into a coherent structure.
The aim is not to create mathematically tidy clusters. It is to produce groups that make sense for SEO and help you make sound content decisions.
Moving beyond spreadsheets
Spreadsheets have long been central to keyword research, and they remain extremely useful. But their format encourages us to think about keywords as individual rows rather than connected ideas.
Graph theory gives us another way to look at the same data: as a connected system of topics, subtopics and concepts.
When datasets reach tens or hundreds of thousands of keywords, that perspective can reveal patterns that would be extremely difficult to spot manually.
It also opens up possibilities for more sophisticated SEO tools. Clusters can be displayed as interactive networks, relationships between topics can be measured, and algorithms can identify central concepts.
New keywords can also be added to an existing graph and assigned according to their strongest relationships.
At its heart, graph theory provides a useful framework for something keyword research has always tried to understand: how searches relate to one another.
Combined with semantic embeddings, community detection and traditional SEO signals, graph-based clustering offers a powerful way to turn large, messy keyword lists into understandable topic structures.
The benefit goes beyond a tidier spreadsheet. It is a more realistic model of the search landscape, one that helps you see both the groups of searches and the connections between them.

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Soft clustering and confidence scores: handling ambiguous keywords honestly

How to cluster 100,000 keywords without comparing every pair
