Choosing a Faiss index for SEO embeddings: exact search, HNSW, IVF and product quantisation
Faiss supports several vector-index families with different speed, recall and memory trade-offs. Learn how to choose an index for keyword and page embeddings without overengineering.

Farky Rafiq
Founder of ClusterIQ

When you are working with thousands of keywords from Ahrefs or Semrush, you eventually move beyond simple keyword matching and into the world of embeddings. However, choosing a vector index is not the same thing as choosing an embedding model. The model determines how your data is represented in a mathematical space, while the index determines how efficiently ClusterIQ can find and retrieve the "nearest neighbours" within that space.
Faiss offers several index families, ranging from perfect, brute-force search to compressed, approximate structures. The right choice for your content plan or URL mapping depends on your corpus size, how fast you need results, and the specific cost of missing a relevant keyword relationship.
Start with exact search as the baseline
Faiss flat indexes store vectors directly and perform an exact search. They provide the gold-standard result against which all other approximate methods should be measured. For many standard SEO tasks, such as mapping 5,000 keywords to a few hundred URLs, exact search is fast enough that you do not need any complex approximate infrastructure at all.
Do not approximate before you need to
If a ClusterIQ workspace contains 20,000 keywords, adding a complex IVF or compressed index can create unnecessary operational overhead without providing a meaningful benefit to the user. Always benchmark the simplest option first. Approximation only becomes compelling when exact search begins to dominate your processing time, memory bandwidth, or the responsiveness of your reporting tools.
HNSW trades memory for fast high-recall search
HNSW (Hierarchical Navigable Small Worlds) builds a navigable graph of vectors. It is a popular choice when you have plenty of memory available and need very low latency for your queries. It is particularly useful when high recall is vital, meaning you cannot afford to miss closely related keywords, and your dataset is large but hasn't yet reached the scale where graph overhead becomes unmanageable. ClusterIQ's HNSW guide covers the specific graph parameters and recall testing in more detail.
IVF partitions the search space first
Inverted-file, or IVF, indexes work by training "coarse centroids" and assigning vectors to specific lists. When you run a query, the system only searches a subset of these lists rather than the whole dataset. This scales very effectively for larger SEO projects, but your recall depends on whether the correct neighbours actually sit inside the lists you chose to probe. If you don't search enough lists, you might miss the very keyword cluster you were looking for.
Product quantisation reduces memory
Product quantisation (PQ) compresses vectors into compact codes. This allows much larger datasets to fit into memory and can speed up distance computations. The trade-off is approximation error caused by the compression itself. For ClusterIQ, PQ becomes relevant when you are dealing with millions of page or keyword vectors that create a genuine memory constraint, rather than using it just because compression sounds sophisticated.
Indexes can be combined
Faiss allows you to combine these methods, such as using IVF with product quantisation. In this setup, a coarse partition reduces the search region, and compressed codes reduce the memory footprint within that region. However, be mindful that each added layer introduces another parameter to tune and another potential reason for missing a relevant neighbour.
Worked example: three corpus sizes
25,000 vectors: If you are building a content brief for a niche site, exact flat search is likely entirely adequate and perfectly accurate.
500,000 vectors: For a large e-commerce site audit, HNSW or IVF may materially improve the speed of repeated neighbour retrieval during clustering.
20 million vectors: At this enterprise scale, the memory and search costs will almost certainly justify using IVF, quantisation, or even distributed indexing.
Corpus size is a major factor, but your hardware, the volume of queries, and your specific recall requirements also dictate the best path forward.
Measure recall against the flat index
To ensure your SEO data remains reliable, you should perform a frozen benchmark:
- Retrieve the exact top-k neighbours using a flat index.
- Retrieve the same from your chosen approximate index.
- Calculate the recall@k (how many the approximate index got right).
- Inspect any important misses to see if they change your clustering outcome.
- Measure the actual latency and memory usage.
Use different indexes for different jobs if necessary
Different SEO tasks have different requirements. An interactive UI search for a single keyword might prioritise low latency so the tool feels snappy. Conversely, an overnight clustering run for a massive site migration might prioritise maximum recall to ensure no pages are left behind. ClusterIQ does not need one single vector index configuration to serve every feature; you can pick the right tool for the specific job.
Index metric must match the embedding setup
Cosine similarity is commonly used in SEO embeddings, often implemented by normalising vectors and using an inner product. You must ensure your index metric matches your embedding model. Do not switch the index metric without revalidating your neighbour order and similarity thresholds, or your clusters will lose their semantic meaning.
Index training data matters for IVF and PQ
Structures that require training, like IVF, should use representative vectors. If your training sample is dominated by one specific market or content type, the resulting partitions might work poorly for other topics. For international ClusterIQ datasets, ensure you sample across different languages and topic families deliberately to keep the index balanced.
Updates and deletions affect operations
Some index types handle new data more naturally than others. If your ClusterIQ project involves frequently adding new page embeddings as content is published, the operational cost of rebuilding or compacting the index should be a key part of your architecture decision.
Keep index version in the run manifest
If your keyword groupings change because you tweaked the index parameters, you need to be able to distinguish that from a change in the underlying model. Always record:
- The index family and metric.
- The version of the training sample used.
- Specific search parameters.
- The build date and the recall benchmark results.
Do not expose infrastructure choices as product claims
Clients and stakeholders care whether ClusterIQ returns reliable, actionable relationships quickly. They generally do not need to hear about HNSW or IVF as evidence of SEO quality. These index choices belong in technical documentation and diagnostics, not as a marketing shortcut.
Practitioner principle: use the simplest index that meets your corpus, latency, and recall requirements, and always validate your approximations against an exact search baseline.
ClusterIQ Conclusion
Faiss provides ClusterIQ with a versatile toolkit to scale embedding retrieval, from perfect flat searches to sophisticated graph and compressed indexes. The right choice is always empirical. You should benchmark recall, latency, and memory on your actual keyword workloads rather than simply adopting the most advanced-sounding index by default.
How this fits the wider ClusterIQ workflow
This technical indexing stage is just one part of a larger evidence chain. ClusterIQ begins with clean source data, preserves lexical features, and creates semantic relationships. By testing the resulting structure before connecting groups to internal links or roadmap actions, we ensure the final recommendations are robust. Keeping these stages separate makes the final SEO strategy easier to explain to clients and safer to adjust as the project evolves.
The system is designed to provide enough detail for an SEO practitioner to challenge a result without needing a degree in data science. A clear explanation can highlight the strongest evidence for a cluster and the practical consequence of an action. Once implemented, these topics can be monitored via Search Console or crawl data. This closes the loop, allowing ClusterIQ to record whether the approved structural changes actually improved search ownership and visibility over time.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Approximate nearest-neighbour recall for SEO: measuring what the fast index misses

Cosine, dot product or Euclidean distance: choosing a similarity measure for keyword embeddings
