Louvain vs Leiden for keyword clustering: community detection compared
Louvain and Leiden organise keyword graphs into communities, but they have different structural properties. Learn what changes, what does not and how to evaluate them for SEO.

Farky Rafiq
Founder of ClusterIQ

You have 2,000 keywords and want to find useful topic groups before deciding what content to create. Louvain and Leiden can help, but they do not read the queries and decide which ones belong on the same page. They work with a network of relationships you have already built.
In that network, called a keyword graph, each keyword is a node and each selected relationship is an edge. Louvain and Leiden are community-detection algorithms: they look for groups within that structure.
This distinction matters. If weak semantic similarities have created links between terms that do not belong together for your SEO task, neither algorithm can repair that judgement for you. Comparing the two means looking at both how they work and how much confidence you should place in the groups they return.
What community detection is doing
A keyword graph might contain a few thousand keywords and many connections, with weights indicating how strong each connection is. Community detection looks for parts of the graph where keywords are more strongly connected to one another than to the rest of the network.
For SEO, you can treat these groups as candidate topical communities. They can help you spot broad subjects, subtopics, terms that bridge different topics and neighbouring areas of meaning.
They are not automatically groups of keywords for individual pages. One community can contain several search intents or page types, so you still need to decide what the grouping means for your content plan.
How Louvain works at a high level
Louvain uses a practical search method, known as a heuristic, to optimise a measure called modularity. Modularity assesses how strongly a network is divided into communities, compared with the connections expected under a reference model.
As the NetworkX documentation explains, Louvain repeatedly moves nodes between communities when doing so improves modularity. It then combines the communities into a smaller network and repeats the process.
This multi-level approach helps Louvain work efficiently on large graphs and has made it popular for exploratory network analysis.
In a keyword graph, edge weights might represent semantic similarity or a score combining several kinds of relationship. Louvain uses those connections to find a partition, meaning a division of the graph, with relatively dense connections within communities and sparser connections between them.
What the resolution parameter changes
Louvain implementations commonly include a resolution setting. In NetworkX, values below one tend to favour larger communities, while values above one tend to favour smaller communities.
For keyword clustering, this setting influences how broad or detailed your topic groups become.
At a lower resolution, you might get one broad “running shoes” community. At a higher resolution, that could split into stability shoes, trail running shoes, racing shoes and beginner running shoes.
Neither result is inherently correct. A broad grouping might help you explore a market, while a more detailed grouping might be more useful when considering which keywords could map to particular URLs.
Why Leiden was developed
Leiden was developed to address a structural problem with Louvain. Traag, Waltman and van Eck showed that Louvain can return communities that are poorly connected and, in some cases, disconnected. A disconnected community contains parts with no path linking them within that community.
Leiden adds a refinement phase and, under the algorithm's conditions, guarantees that the communities it identifies are connected. In the authors' experiments, it also produced higher-quality partitions and ran faster than Louvain on the networks they tested.
That does not mean Leiden is automatically “more accurate” at understanding search intent. The practical benefit is a stronger guarantee about how the keywords within each graph community are connected.
Connectivity is not the same as semantic correctness
A connected group can still be unsuitable for the SEO decision you need to make.
Imagine a network containing these queries:
- “best accounting software”;
- “accounting software pricing”;
- “free accounting spreadsheet”;
- “what is accounting software”.
Connections based on meaning may bring these terms into a coherent region around accounting software. But the page formats needed to serve those searches still differ: a comparison, pricing information, a spreadsheet resource and an explanation are not interchangeable.
Leiden improves the network partition according to its graph objective. It does not add knowledge about commercial intent, page type or what the searcher wants to do unless you have already encoded those signals in the graph.
Graph construction usually matters more than switching algorithms
Before comparing Louvain and Leiden, check the graph they will be working with.
The important choices include:
- how keywords are represented when finding potential neighbours, such as with text features or semantic vectors;
- which similarity measure you use;
- which similarity threshold or nearest-neighbour rule decides whether keywords get connected;
- whether connections have weights;
- whether you combine lexical evidence, meaning similarities in wording, with semantic and search engine results page (SERP) evidence;
- whether you retain weak connections that bridge otherwise separate groups.
Our guide to choosing a similarity threshold explains how a seemingly small change to the connection rule can radically change graph density, or how many connections the graph contains.
If the graph links unrelated terms, a more sophisticated community-detection algorithm will still be organising the wrong relationships.
Weighted edges are useful, but they encode assumptions
Community-detection algorithms can work with weighted graphs. This is useful because some keyword relationships are stronger than others.
For example, a semantic similarity score of 0.91 might carry more weight than 0.76. You might also strengthen a connection when similarity in meaning and wording agree, or when two queries share meaningful overlap in their search results.
The weights still need to be understandable. If you combine several signals into one score without recording how each was scaled, the graph becomes difficult to reproduce and harder to troubleshoot.
A good production system stores the component scores as well as the final edge weight. That gives you a way to investigate why a particular connection was strong enough to influence a group.
Randomness and reproducibility
Louvain results can depend on the order in which nodes are considered and on random choices made during the process. NetworkX provides a seed parameter to make those choices reproducible. Leiden implementations also commonly use randomisation during optimisation.
For an SEO workflow, save the seed alongside:
- the graph-construction configuration;
- the algorithm and version;
- the resolution;
- the edge-weight definition;
- the source dataset version.
Keeping these details does not prove that the communities are objectively correct. It gives you a reproducible setup so you can rerun the analysis and make meaningful comparisons.
Test community stability
A useful topic group should not depend entirely on one fragile combination of settings.
Try nearby resolution values and, where relevant, several seeds. Check which groups stay together, which consistently split into smaller groups and which individual queries move between communities.
A query that moves frequently is not necessarily “bad data”. It may genuinely bridge two topics, or it may show that the boundary between them is weak.
This is one benefit of looking at keywords as a graph. An unstable grouping can itself point to a relationship worth reviewing, rather than being something to dismiss.
When Louvain remains useful
Louvain remains useful for fast exploratory analysis, particularly if your existing tools already support it and you understand its limitations.
It can provide a strong baseline for:
- testing the rules used to build your graph;
- exploring different resolution levels;
- visualising candidate topical communities;
- comparing changes between datasets.
These are reasonable uses when the output informs a decision rather than making it automatically, and an SEO practitioner reviews the results.
When Leiden is attractive
Leiden is worth considering when the structural quality of your communities matters and you want the stronger connectivity guarantees described in the original paper.
It is particularly worth evaluating if Louvain produces:
- large communities held together by fragile links;
- unexpected disconnected subgroups;
- unstable partitions after repeated optimisation;
- communities that look structurally weak when visualised.
Even then, test whether switching to Leiden makes the output more useful for your SEO task. Better graph structure should not be treated as proof of better content recommendations.
A practical comparison workflow
- Freeze the graph. Give both algorithms exactly the same nodes, edges and weights so you are comparing the algorithms, not different inputs.
- Run Louvain and Leiden at comparable granularity. Compare groups at a similar level of detail. Do not mistake the effect of a resolution change for an algorithm improvement.
- Inspect connectivity. Look inside communities for disconnected regions or groups linked only weakly to one another.
- Compare stability. Repeat the analysis with controlled seeds and nearby resolution values to see how much the groupings change.
- Review bridge terms. Identify keywords that frequently move between communities and consider whether they genuinely connect different topics.
- Validate against the task. Check search intent, page type and business relevance before turning communities into proposed URLs or categories.
ClusterIQ Conclusion
Louvain and Leiden let you explore keyword relationships as a network, rather than treating clustering only as a way to assign each spreadsheet row to a group.
Louvain is an efficient, widely used heuristic for optimising modularity. Leiden addresses structural weaknesses in Louvain and provides stronger guarantees that the communities it identifies are connected.
For SEO, though, the graph is still the model you have built. Community detection organises the relationships you give it; it does not replace judgement about intent or content. The strongest workflow combines careful graph construction, reproducible settings, stability testing and practitioner review.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

The resolution parameter in graph clustering: how one setting changes topic granularity

Graph centrality for SEO: finding hubs, bridges and misleadingly important keywords
