Skip to main content
All articles
Clustering
24 July 2026 4 min read

Cross-encoder reranking for SEO: when a second-stage model improves keyword relationships

Bi-encoders retrieve candidate relationships quickly; cross-encoders can score those pairs more precisely. Learn where reranking helps SEO clustering and where it becomes too expensive.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Editorial diagram showing fast semantic retrieval narrowing many keyword relationships into a few boundary cases for precise cross-encoder reranking before forming cleaner clusters.

When you are staring at a spreadsheet of 5,000 keywords exported from Ahrefs or Search Console, you are essentially trying to solve two different problems. First, you need to find which terms generally belong together. Second, you need to decide exactly where the boundaries lie between different pages or content briefs. In the world of machine learning, these tasks require different tools.

Bi-encoders, like the Sentence Transformers we often use for clustering, are brilliant at the first part. They allow us to quickly compare thousands of keywords to find likely neighbours. However, cross-encoders take a smaller selection of these pairs and look at them much more closely. They produce a far more precise judgement of how two phrases relate, though they require significantly more processing power to do so.

For ClusterIQ, this suggests a two-stage approach rather than relying on a single model for every single calculation.

Why cross-encoders are different

A bi-encoder creates a mathematical map (a vector) for every keyword independently. A cross-encoder, however, processes two keywords simultaneously, allowing the model to pay attention to the specific relationship between the two phrases before giving them a score.

This joint processing picks up on subtle nuances that a standard embedding might miss. The catch is speed. If you have 50,000 keywords, it is simply not practical to run every possible combination through a cross-encoder.

Retrieve first, rerank second

The standard advice from the Sentence Transformers documentation is to follow a retrieve and rerank workflow:

  1. Use a fast retriever or bi-encoder to identify a small group of candidates.
  2. Pass those specific candidates through a cross-encoder.
  3. Rerank or filter the results based on these more accurate scores.

This fits keyword research perfectly. ClusterIQ can find semantic neighbours in seconds, then apply a deeper model only where the distinction really matters for your content plan.

Reranking is most valuable near the boundary

Obvious pairs do not need a second opinion. You do not need an expensive model to tell you that "affordable running shoes" and "cheap running trainers" belong on the same page. The real value of reranking appears when dealing with:

  • Terms with the same topic but different search intents.
  • Products in the same family but with different model numbers.
  • Short, ambiguous queries that could sit in multiple clusters.
  • Keywords that sit right on the edge of a similarity threshold.
  • High-value terms where the clustering decision dictates a major site architecture change.

Cross-encoder scores are not automatically probabilities

It is important to note that different models output different types of scores. The Sentence Transformers documentation points out that some models return raw numbers (logits) that have not been converted into a 0 to 1 scale.

You should avoid telling a client that a pair is "92% likely to be a match" unless the model has been specifically calibrated for that. ClusterIQ treats these outputs as evidence with a specific scale, rather than a simple percentage.

Worked example: semantic similarity disagrees with page purpose

Consider these two keywords:

  • "project management software pricing"
  • "project management software tutorial"

A standard bi-encoder might group these together because the subject matter is identical. However, a cross-encoder trained to spot relevance might better identify that the user is trying to do two different things.

Even so, the final SEO decision should still consider SERP evidence and page types. The cross-encoder is a sharper semantic judge, but it is not a replacement for an SEO strategy.

Use reranking before graph construction or after candidate retrieval

There are two practical ways to use this second stage.

Before building the cluster graph: Rerank the candidate pairs and only keep those that pass the stricter cross-encoder rule.

For manual review: Build your clusters using efficient embeddings first, then only use the cross-encoder to check "weak bridges" or disputed keywords that seem to sit between two different clusters.

The second method is much faster and usually provides everything a search marketer needs.

Pair scoring can support URL matching too

ClusterIQ can also use this to match keywords to existing pages. We can retrieve potential URLs using embeddings and then use the cross-encoder to confirm which page is the best fit. This is particularly useful when you have several pages covering similar ground but one is a much better match for the specific query. This approach works alongside embedding-based URL mapping.

Do not use the reranker to hide a poor first stage

If the bi-encoder fails to find the right keyword in the first place, the cross-encoder will never see it. You should evaluate your retrieval process separately. Check how often the correct "neighbour" keyword appears in your top 20 or 50 results. The second stage is there to improve precision, not to fix a broken retrieval system.

Fine-tuning may help, but it raises the governance burden

The Sentence Transformers documentation notes that these models perform best when tuned on specific data. While ClusterIQ could eventually use human-approved keyword relationships for fine-tuning, this needs to be handled carefully. Human decisions can be inconsistent, and a model is only as good as the logic used to train it.

Cost should follow consequence

You do not need to rerank every single relationship just because you can. A sensible approach is to save the deep processing for:

  • High-volume or high-value "money" terms.
  • Keywords that are hovering right on the edge of a cluster.
  • Terms with mixed signals or multiple meanings.
  • Keywords that act as bridges between two large topics.
  • URL mapping decisions that involve significant developer time to implement.

This keeps the workflow fast while focusing the "AI brainpower" where a mistake would actually cost you money.

Store both stages for explainability

If a keyword pair has a similarity score of 0.85 from the first model, but the cross-encoder lowers that confidence, we should keep both numbers. This allows you to see exactly where the deeper model disagreed with the initial automated grouping, rather than just seeing a final, unexplained result.

Practitioner principle: Use the fast model to find the likely candidates and the expensive model to settle the difficult debates. Do not pay for deep judgement when the answer is already obvious.

ClusterIQ Conclusion

Cross-encoder reranking adds a layer of precision to SEO keyword clustering without making the process too slow or expensive. By using bi-encoders for the heavy lifting and reserving cross-encoders for the high-stakes decisions, ClusterIQ creates a balanced system where the level of effort matches the level of risk.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.