Cross-encoder reranking for SEO: when a second-stage model improves keyword relationships
Bi-encoders retrieve candidate relationships quickly; cross-encoders can score those pairs more precisely. Learn where reranking helps SEO clustering and where it becomes too expensive.

Farky Rafiq
Founder of ClusterIQ

When you are staring at a spreadsheet of 5,000 keywords exported from Ahrefs or Search Console, you are essentially trying to solve two different problems. First, you need to find which terms generally belong together. Second, you need to decide exactly where the boundaries lie between different pages or content briefs. In the world of machine learning, these tasks require different tools.
Bi-encoders, like the Sentence Transformers we often use for clustering, are brilliant at the first part. They allow us to quickly compare thousands of keywords to find likely neighbours. However, cross-encoders take a smaller selection of these pairs and look at them much more closely. They produce a far more precise judgement of how two phrases relate, though they require significantly more processing power to do so.
For ClusterIQ, this suggests a two-stage approach rather than relying on a single model for every single calculation.
Why cross-encoders are different
A bi-encoder creates a mathematical map (a vector) for every keyword independently. A cross-encoder, however, processes two keywords simultaneously, allowing the model to pay attention to the specific relationship between the two phrases before giving them a score.
This joint processing picks up on subtle nuances that a standard embedding might miss. The catch is speed. If you have 50,000 keywords, it is simply not practical to run every possible combination through a cross-encoder.
Retrieve first, rerank second
The standard advice from the Sentence Transformers documentation is to follow a retrieve and rerank workflow:
- Use a fast retriever or bi-encoder to identify a small group of candidates.
- Pass those specific candidates through a cross-encoder.
- Rerank or filter the results based on these more accurate scores.
This fits keyword research perfectly. ClusterIQ can find semantic neighbours in seconds, then apply a deeper model only where the distinction really matters for your content plan.
Reranking is most valuable near the boundary
Obvious pairs do not need a second opinion. You do not need an expensive model to tell you that "affordable running shoes" and "cheap running trainers" belong on the same page. The real value of reranking appears when dealing with:
- Terms with the same topic but different search intents.
- Products in the same family but with different model numbers.
- Short, ambiguous queries that could sit in multiple clusters.
- Keywords that sit right on the edge of a similarity threshold.
- High-value terms where the clustering decision dictates a major site architecture change.
Cross-encoder scores are not automatically probabilities
It is important to note that different models output different types of scores. The Sentence Transformers documentation points out that some models return raw numbers (logits) that have not been converted into a 0 to 1 scale.
You should avoid telling a client that a pair is "92% likely to be a match" unless the model has been specifically calibrated for that. ClusterIQ treats these outputs as evidence with a specific scale, rather than a simple percentage.
Worked example: semantic similarity disagrees with page purpose
Consider these two keywords:
- "project management software pricing"
- "project management software tutorial"
A standard bi-encoder might group these together because the subject matter is identical. However, a cross-encoder trained to spot relevance might better identify that the user is trying to do two different things.
Even so, the final SEO decision should still consider SERP evidence and page types. The cross-encoder is a sharper semantic judge, but it is not a replacement for an SEO strategy.
Use reranking before graph construction or after candidate retrieval
There are two practical ways to use this second stage.
Before building the cluster graph: Rerank the candidate pairs and only keep those that pass the stricter cross-encoder rule.
For manual review: Build your clusters using efficient embeddings first, then only use the cross-encoder to check "weak bridges" or disputed keywords that seem to sit between two different clusters.
The second method is much faster and usually provides everything a search marketer needs.
Pair scoring can support URL matching too
ClusterIQ can also use this to match keywords to existing pages. We can retrieve potential URLs using embeddings and then use the cross-encoder to confirm which page is the best fit. This is particularly useful when you have several pages covering similar ground but one is a much better match for the specific query. This approach works alongside embedding-based URL mapping.
Do not use the reranker to hide a poor first stage
If the bi-encoder fails to find the right keyword in the first place, the cross-encoder will never see it. You should evaluate your retrieval process separately. Check how often the correct "neighbour" keyword appears in your top 20 or 50 results. The second stage is there to improve precision, not to fix a broken retrieval system.
Fine-tuning may help, but it raises the governance burden
The Sentence Transformers documentation notes that these models perform best when tuned on specific data. While ClusterIQ could eventually use human-approved keyword relationships for fine-tuning, this needs to be handled carefully. Human decisions can be inconsistent, and a model is only as good as the logic used to train it.
Cost should follow consequence
You do not need to rerank every single relationship just because you can. A sensible approach is to save the deep processing for:
- High-volume or high-value "money" terms.
- Keywords that are hovering right on the edge of a cluster.
- Terms with mixed signals or multiple meanings.
- Keywords that act as bridges between two large topics.
- URL mapping decisions that involve significant developer time to implement.
This keeps the workflow fast while focusing the "AI brainpower" where a mistake would actually cost you money.
Store both stages for explainability
If a keyword pair has a similarity score of 0.85 from the first model, but the cross-encoder lowers that confidence, we should keep both numbers. This allows you to see exactly where the deeper model disagreed with the initial automated grouping, rather than just seeing a final, unexplained result.
Practitioner principle: Use the fast model to find the likely candidates and the expensive model to settle the difficult debates. Do not pay for deep judgement when the answer is already obvious.
ClusterIQ Conclusion
Cross-encoder reranking adds a layer of precision to SEO keyword clustering without making the process too slow or expensive. By using bi-encoders for the heavy lifting and reserving cross-encoders for the high-stakes decisions, ClusterIQ creates a balanced system where the level of effort matches the level of risk.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Choosing an embedding model for keyword clustering: accuracy is only one part of the decision

Sentence Transformers for SEO: what bi-encoders actually add to keyword analysis
