Skip to main content
All articles
Clustering
25 July 2026 4 min read

Sentence Transformers for SEO: what bi-encoders actually add to keyword analysis

Sentence Transformers make semantic comparison practical at scale by encoding text independently. Learn what that enables for SEO and where other evidence is still needed.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing keywords encoded independently into reusable vector nodes, then connected as a sparse graph of candidate semantic neighbours with SEO evidence filtering the final relationships.

Sentence Transformers have fundamentally changed how we handle semantic text comparison in SEO. If you have ever tried to manually group five thousand keywords from Ahrefs or Search Console, you know the struggle: words that look different often mean the same thing, while identical looking phrases can have completely different intents.

In the past, comparing every keyword to every other keyword was computationally expensive. Sentence Transformers solve this by using a bi-encoder approach. Instead of looking at two phrases together, the model creates a single, reusable mathematical map (a vector) for each keyword. These maps can be stored and compared instantly, making it possible for ClusterIQ to process large datasets without needing a massive server for every calculation.

What a bi-encoder does

Think of a Sentence Transformer as a translator that turns text into a fixed set of coordinates. Once a keyword has its coordinates, we can use them for various tasks: finding similar terms, grouping keywords into clusters, or matching new search queries to existing landing pages. Because we only have to translate the keyword once, we can reuse that information across multiple SEO reports.

Why pairwise transformer scoring is different

There is another method called a cross-encoder, which reads two pieces of text simultaneously to make a very precise judgement. While accurate, it is incredibly slow because it has to re-evaluate every possible pair. If you are dealing with ten thousand keywords, checking every pair would require millions of individual calculations. By using a bi-encoder, ClusterIQ encodes the list once and then quickly finds the most likely neighbours, saving hours of processing time.

Semantic relationships go beyond exact wording

The real magic of embeddings is their ability to connect synonyms and paraphrases even when they share no actual words. A search for "cheap trainers" and "affordable running shoes" might not share any vocabulary, but the model understands they belong together. This is a significant step up from older methods, as we explored in our guide on TF-IDF versus semantic embeddings.

The embedding still reflects its training

It is important to remember that these models do not inherently understand SEO. A general model might recognise that "iPhone 15 price" and "iPhone 15 reviews" are related to the same phone, but it might not realise that an SEO needs to keep these on separate pages due to different user intents. This is why ClusterIQ treats semantic similarity as just one piece of the puzzle rather than the final word on how to group your content.

Model-specific similarity functions matter

Not all models measure "distance" the same way. Some prefer cosine similarity, while others use Euclidean distance or dot products. Choosing the wrong metric can lead to messy clusters. It is always best to check the specific model documentation. For a deeper dive into why this matters, see our metric comparison, which explains how normalising your data changes the results.

Batch encoding matters in production

When you are building a content plan from a massive Semrush export, efficiency is everything. Processing keywords one by one is a waste of resources. A professional workflow involves deduplicating the list, encoding them in batches, and caching the results. This allows ClusterIQ to tweak clustering rules or thresholds without having to start the whole mathematical process from scratch.

Sentence embeddings enable approximate search

For very large projects, we use systems like Faiss to perform "approximate" searches. Instead of checking every single keyword, the system quickly identifies the most likely candidates. This technique is vital for the architecture we use when clustering 100,000 keywords efficiently.

Short queries can be difficult

SEO data is often messy. A two-word keyword like "Apple support" could refer to the fruit or the tech giant. Because short queries lack context, the embedding alone might struggle. We find that adding external data, such as entity information or existing ranking URLs, helps clear up this confusion.

Exact modifiers need protection

Sometimes, the model is too smart for its own good. It might see "blue widget size M" and "blue widget size L" as identical because the surrounding words are the same. For technical SEO and ecommerce, ClusterIQ combines these semantic vectors with exact word matching to ensure important modifiers like sizes, locations, or model numbers are not lost in the shuffle.

Worked example: candidate generation

If you upload 50,000 keywords to ClusterIQ, the practical workflow looks like this:

  1. Clean and deduplicate the list to save time.
  2. Turn each unique query into a vector using a Sentence Transformer.
  3. Quickly find the most likely semantic neighbours for each term.
  4. Filter out weak matches that don't make sense for SEO.
  5. Create a map of how these keywords connect.
  6. Run an analysis to find the natural "communities" or clusters.
  7. Present the final groups to the user for review.

The model suggests the relationships, but the final page structure is still guided by SEO logic.

Use cross-encoders where extra precision is worth it

While bi-encoders are great for the heavy lifting, cross-encoders are excellent for double-checking difficult cases. You can use the fast bi-encoder to find a hundred potential matches and then use a cross-encoder to pick the absolute best one. This hybrid approach is perfect for high-value keywords where accuracy is non-negotiable.

Reproducibility still matters

Always keep track of which model version you used. If you update the model mid-project, your clusters might shift. By storing the model details and the vectors together, you can ensure your reports remain consistent and verify if a new model actually provides better SEO insights or just different ones.

Use human benchmarks alongside model benchmarks

Standard AI benchmarks are fine, but they don't always reflect the real world of search. We recommend testing models against a set of "gold standard" SEO examples: terms with clear intent shifts, local variations, and tricky product names. This keeps the methodology grounded in what actually works for clients and business owners.

Practitioner principle: Sentence Transformers make semantic relationships cheap enough to use at scale. They do not make those relationships automatically correct for SEO.

ClusterIQ Conclusion

Sentence Transformers provide the scalable foundation ClusterIQ needs to handle thousands of keywords at once. They offer a fast, reusable way to understand how search terms relate to one another.

However, their true value comes when they are paired with human expertise and real-world data like intent and existing URLs. The AI expands your view of the data, but the marketer still decides how to turn those relationships into a winning content strategy.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.