Approximate nearest-neighbour recall for SEO: measuring what the fast index misses
Fast vector search is useful only if it retrieves the neighbours the downstream SEO workflow needs. Learn how recall@k and error analysis keep approximation honest.

Farky Rafiq
Founder of ClusterIQ

You have 2,000 keywords from Ahrefs, Semrush or Search Console and need a content plan, not hours of processing. Approximate nearest-neighbour search, or ANN, can make finding related keyword vectors dramatically faster. Vectors are numerical representations of those keywords.
But speed needs a quality check: does the fast index retrieve enough of the neighbours that exact search would find for the SEO task?
Recall@k is the basic measure
For each test query, retrieve the exact top-k neighbours and the approximate top-k neighbours. Here, k is simply the number requested.
Recall@k is the proportion of exact neighbours recovered. If approximate search finds 19 of the exact top 20, recall@20 is 0.95.
Average recall can hide difficult queries
Average recall of 0.98 looks excellent. It can still hide queries scoring 0.60, particularly those involving:
- rare entities;
- product model codes;
- another language;
- unusual vector norms, meaning vector lengths;
- sparse topic regions with few nearby vectors.
ClusterIQ should report the spread of recall scores and recurring failure types, not just the mean.
Worked example: missing the one neighbour that matters
A keyword has 20 strong semantic neighbours. ANN retrieves 19 but misses the only query sharing the exact product model.
The score remains high, yet that missing relationship could change an entity-aware graph and the page assigned to the keyword. Review the consequences of misses, not just their number.
Benchmark against exact search on a frozen sample
Exact retrieval provides the reference. Keep a fixed, representative benchmark containing:
- head terms;
- long-tail queries;
- rare entities;
- several languages;
- commercial and informational intents;
- page embeddings as well as keyword embeddings.
A frozen sample makes configuration comparisons meaningful.
Recall and latency form a trade-off curve
Most ANN indexes have settings controlling search effort. Increasing effort usually improves recall but also increases latency, the time a search takes.
Plot both rather than copying a setting from documentation. Look for where extra waiting time stops buying meaningful recall for the SEO workflow.
Measure at the candidate-set size you actually use
If ClusterIQ retrieves 50 candidates before reranking them, recall@10 alone is insufficient. Measure recall@50 and whether the correct page or neighbour appears anywhere in that set.
A second-stage model can reorder candidates. It cannot recover a useful URL that retrieval never supplied.
Retrieval recall differs from semantic relevance
Exact neighbours are the reference for testing the index, not a guarantee of good SEO relationships. Keep two evaluation layers separate:
- Index recall: did ANN reproduce exact vector search?
- Task quality: did vector search retrieve relationships an SEO would find useful?
Perfect index recall can still leave you with unsuitable groupings for a content brief.
Model changes reset the benchmark
A new embedding model changes the vector space and therefore the exact neighbours. After an upgrade, rebuild the exact benchmark and revalidate the ANN index. Previous results no longer establish its recall.
Corpus growth can change performance
An index tuned at 100,000 vectors may behave differently at five million. Track recall as the collection grows and after major bulk imports, rather than assuming the original settings remain suitable.
Index build quality matters
IVF is an index type that needs training. A poor training sample can lower recall in underrepresented topic regions.
HNSW uses a graph, whose construction settings affect connectivity and recall. Search-time tuning is therefore only part of the picture.
Faiss index selection provides the architectural context.
Use hard queries deliberately
Include tightly packed groups of neighbouring vectors and cases where the correct neighbour is only slightly closer than alternatives. These difficult boundaries expose whether speed optimisations are cutting search effort too aggressively.
Map retrieval failures to downstream effects
A miss may be harmless if it changes neither a cluster connection nor a page candidate. Removing the only correct URL from a migration mapping is much more serious.
ClusterIQ could connect retrieval QA to the workflows using those candidates, making that distinction visible.
Choose settings by workload
Offline clustering can tolerate longer searches for better recall. Interactive exploration may accept slightly lower recall to stay responsive.
High-value automated actions should use more conservative settings or an exact final check. A quick exploratory view and a consequential URL decision need not share one configuration.
Record benchmark evidence with the index version
For every production index version, save:
- corpus size;
- index type;
- parameters;
- recall@k;
- latency distribution;
- benchmark sample;
- known failure cases.
This gives reviewers evidence they can revisit, rather than an isolated score.
Practitioner principle: approximate search earns its place by preserving the neighbours the workflow needs at a lower cost. Measure both parts of that trade-off.
ClusterIQ Conclusion
Recall@k gives ClusterIQ an objective way to check fast vector search. Benchmarking against exact neighbours, then reviewing high-consequence misses, would help retrieval scale without quietly weakening the evidence behind clustering and URL mapping.
How ClusterIQ would validate this before production use
One plausible output is not enough. ClusterIQ would run the current and proposed configurations against a frozen set of representative queries and pages, then compare decisions that changed. Include obvious examples, difficult boundaries, important entities and high-value page mappings.
Keep acceptance evidence layered: the underlying score or feature, the resulting neighbours or cluster membership, and the downstream SEO action. Better internal metrics do not justify moving important queries to less sensible pages. Conversely, a slightly weaker generic metric may be acceptable if it fixes a recurring practitioner error without regressions elsewhere.
Preserving before-and-after examples with the run version would let future changes face the same checks. That creates a regression-tested system, rather than settings that merely look good on the latest dataset.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Choosing a Faiss index for SEO embeddings: exact search, HNSW, IVF and product quantisation

Cosine, dot product or Euclidean distance: choosing a similarity measure for keyword embeddings
