HNSW for SEO embeddings: fast neighbour search without pretending approximate means exact
HNSW builds a navigable graph for approximate nearest-neighbour search. Learn how recall, speed and graph parameters affect large-scale keyword and page matching.

Farky Rafiq
Founder of ClusterIQ

When you are managing a few thousand keywords from Ahrefs or Search Console, comparing every keyword to every other keyword is a breeze for modern computers. However, as your data grows to hundreds of thousands or even millions of rows, the maths changes. If you try to find the "nearest neighbours" for every single keyword in a massive list using exact calculations, your processing time and costs will skyrocket.
Hierarchical Navigable Small World graphs, or HNSW, are the industry standard for solving this. It is a method for "approximate" nearest-neighbour search. Instead of checking every possible match, it builds a clever multi-layer map that allows the system to zoom into the right neighbourhood almost instantly.
For ClusterIQ, HNSW acts as a high-performance engine. It makes large-scale analysis possible, but we have to ensure that "approximate" does not mean we lose the vital connections that make your content clusters accurate.
What HNSW is solving
Imagine you have one specific keyword and you need to find the most similar terms from a library of a million others. An exact search would look at every single one of those million items. It is perfectly accurate but incredibly slow at scale.
Approximate search makes a deal: it trades a tiny, controlled chance of missing the absolute best match for a massive increase in speed. When we are building content plans or mapping URLs, this trade-off is usually a huge win, provided the retrieval stage is followed by more precise checks.
Why the graph is hierarchical
HNSW works by creating layers of connections. Think of it like a motorway network. The top layers have very few junctions and allow you to travel long distances across the data space quickly. As you get closer to your destination, you move down into the lower layers, which have dense, local connections like residential streets. This allows the search to refine the results until it finds the best candidates around your query.
Approximate does not mean random
It is a mistake to think that approximate search is "guesswork". A properly tuned HNSW index can achieve incredibly high accuracy. The secret is not just using the default settings, but measuring "recall" against the specific types of keywords and entities you are actually processing.
Recall is the quality metric that matters first
To know if the system is working, we use a benchmark. We take a sample of queries and find their matches using the slow, exact method. Then we run the same queries through HNSW and compare the results.
We measure "Recall@k", which basically asks: "Out of the top 20 matches we know exist, how many did the fast system actually find?" If HNSW finds 19 or 20 of them every time, the speed boost is worth it. If it starts missing important variations of your core entities, we know the settings need to be tightened.
Worked example: one million keyword embeddings
Let us say ClusterIQ is processing a million vectors. Instead of comparing a new keyword against all million, the HNSW index quickly grabs the 30 most likely candidates. These 30 then go through a much stricter set of checks for intent, entity matching, and lexical similarity.
In this workflow, the approximate search is not the final judge of which keywords belong in a cluster. It simply acts as a filter, narrowing down a massive field of possibilities to a manageable shortlist for the expensive, high-quality logic to handle.
Graph construction parameters affect memory and quality
When setting up HNSW, there are dials we can turn. We can choose how many connections each point in the graph keeps and how hard the system works to build the map in the first place. More connections usually mean better accuracy, but they also require more memory. We balance these settings based on the size of the keyword set to ensure the tool remains both fast and precise.
Search effort can be tuned separately
One of the best features of HNSW is that we can change how hard it looks for matches at the moment you ask the question. We can use different "effort" levels depending on what you are doing:
- Quick interactive exploration where speed is king.
- Deep offline clustering for a final content strategy.
- High-stakes URL mapping for site migrations.
- Processing massive bulk imports of historical data.
Do not benchmark only on average latency
Speed is not the only factor. When we evaluate performance, we look at a full range of metrics:
- Recall@k (accuracy).
- P50 and P95 latency (how long the typical and the slowest searches take).
- Memory usage.
- How long it takes to build the index.
- How well it handles rare or niche keywords.
- How performance holds up as the keyword list grows.
Often, the "average" search looks fine, but the difficult, niche queries are where the real insights are found.
Embedding normalisation still matters
The way we measure distance between keywords must stay consistent. If we are using cosine similarity to understand how closely related two topics are, we need to ensure our vectors are prepared correctly so the search index reflects that logic. You can read more about this in our guide on Cosine, dot product and Euclidean distance.
HNSW is a retrieval layer, not a clustering algorithm
It is tempting to call any fast search "AI clustering", but that is not quite right. HNSW is just the librarian who finds the books you might want. The actual clustering, which involves community detection and intent modelling, is what decides how those keywords are grouped into a brief. This distinction is vital for the architecture we discuss in semantic search versus clustering.
Index updates need testing
As you add new pages or fresh keywords from Search Console, the HNSW graph evolves. We have to monitor whether these additions make the search slower or less accurate over time. If we make a major change to our underlying AI models, we rebuild the index from scratch to ensure everything stays compatible.
Use exact search as the regression oracle
We do not need to use the slow, exact search for every day-to-day task. However, we keep it as a "source of truth". By comparing HNSW results against this gold standard, we can be sure that any updates to ClusterIQ are actually making the tool better without sacrificing the quality of your clusters.
Keep retrieval errors separate from clustering errors
If a keyword is missing from a cluster, we need to know why. Was it because the search engine never found it (a retrieval error), or because the clustering logic decided it did not belong there? By logging these separately, we can fix the right part of the pipeline and give you better results faster.
Practitioner principle: approximate neighbour search is an engineering trade-off. Measure what it misses before trusting the speed it gains.
ClusterIQ Conclusion
HNSW makes it possible to perform sophisticated semantic analysis on a massive scale without the wait times. By treating it as a high-speed retrieval layer rather than the final word on clustering, we can provide the speed that smart marketers need while maintaining the technical depth that SEO requires. The key is constant measurement against exact results to ensure that your content plans are built on solid ground.
Failure modes are part of the feature, not an appendix
A truly useful tool should tell you where it might be struggling. We look closely at boundary cases, such as very short phrases, rare technical terms, or multilingual data. If the evidence for a keyword relationship is weak, it is often better to show that uncertainty rather than forcing a keyword into a group where it does not fit. This might mean leaving a query unassigned or asking for a human review, which is always safer than providing a "precise" answer that is actually wrong.
By categorising these failures, we can improve the system more effectively. If a keyword never makes it into the candidate list, no amount of tweaking the clustering algorithm will fix it. We ensure that improvements are targeted at the right layer of the process, whether that is the initial data cleaning or the final grouping logic.
Design separate retrieval profiles for separate ClusterIQ jobs
Not every SEO task requires the same level of intensity. A quick look at related queries can be near-instant, while a full-scale site migration deserves a much more thorough search. We use different profiles for different jobs, ensuring that we only spend the extra processing power when the consequences of the decision are high.
This approach also helps us track quality. If you see a relationship in one part of the tool but not another, we can quickly identify if that is due to the speed settings or the semantic model itself. It keeps the performance proportional to the task at hand.
Plan for index lifecycle, not only index creation
A search index is a living thing. As your data grows and changes, we monitor the health of the HNSW graph. We track memory, latency, and accuracy, triggering updates whenever the index starts to drift. This operational rigour ensures that the infrastructure behind your keyword research is just as reliable as the insights you draw from it.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Choosing a Faiss index for SEO embeddings: exact search, HNSW, IVF and product quantisation

Approximate nearest-neighbour recall for SEO: measuring what the fast index misses
