Skip to main content
All articles
Clustering
21 August 2026 4 min read

Weighted keyword graphs: combining semantic similarity, entities and SERP evidence

One relationship score can hide several different kinds of evidence. Learn how to combine semantic, lexical, entity and SERP signals without creating an opaque black box.

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

Two keyword nodes connected by separate semantic, lexical, entity and SERP evidence lines that combine into one weighted graph edge, with a distinct constraint gate showing that critical incompatibilities are checked separately.

A keyword graph does not have to rely on one signal alone. Semantic similarity can catch paraphrases. Lexical similarity can preserve wording that actually matters. Entity overlap can protect product and brand distinctions. SERP overlap can add evidence straight from the current search results.

Combine those signals well and you get a stronger graph. Combine them carelessly and you get a score nobody can explain.

Keep the signals separate before you combine them

For every candidate relationship, store the components individually:

  • semantic similarity;
  • lexical similarity;
  • entity agreement;
  • SERP overlap;
  • intent compatibility;
  • existing-page evidence.

Even once the system produces one final edge weight, these components should stay inspectable behind it.

Different signals are answering different questions

Semantic similarity asks whether the language is close under an embedding model. Lexical similarity asks whether the surface wording overlaps. Entity agreement asks whether the queries refer to the same named objects or attributes. SERP overlap asks whether the current search system retrieves similar URLs for both. Intent compatibility asks whether the likely tasks behind each query can reasonably sit together. Squashing all of that into one number does not make these questions the same question.

Normalise before you combine anything

If one signal ranges from 0 to 1 and another from 0 to 100, adding them directly just lets the larger scale dominate. Any weighted combination needs explicit scaling. Options include:

  • min-max normalisation;
  • z-score standardisation;
  • rank-based transformation;
  • binary constraints for critical evidence.

Record whichever choice you make, because it changes what the final weight actually means.

Some signals should be constraints, not weights

A critical product-model mismatch may deserve a hard rule rather than a small penalty. Examples include:

  • different country markets;
  • different regulated product classes;
  • different product IDs;
  • incompatible page types.

Trying to encode everything as "semantic 0.82 minus entity penalty 0.06" can quietly make an important distinction disappear inside the arithmetic.

Weighting should reflect the job the graph is doing

For broad topic discovery, semantic similarity might deserve most of the weight. For URL mapping, SERP and page-type evidence probably matter more. For ecommerce taxonomy, entity and attribute compatibility might dominate. One universal edge formula is unlikely to work equally well for every part of the product.

Run ablation tests to see what each signal is doing

A powerful way to understand a multi-signal model is to remove one signal at a time and compare:

  • semantic only;
  • semantic plus entities;
  • semantic plus SERP;
  • semantic plus entities plus SERP;
  • the full model including intent.

Check which relationships and clusters actually change. If a signal barely moves anything, it might be unnecessary. If it changes everything, confirm that the effect is one you actually want.

Let the interface explain each edge

If two keywords are connected, a practitioner should be able to see why. For example:

  • semantic similarity: 0.86;
  • shared entity: "HDBSCAN";
  • SERP overlap: 6/10 URLs;
  • intent: informational / informational;
  • final edge: strong.

That is far more useful than a single unexplained score of 0.79.

Use confidence bands rather than fake precision

If your weighted score is a hand-designed combination rather than a calibrated probability, do not present it as if 0.812 means something meaningfully different from 0.807. Operational bands such as strong, moderate, weak and review-required can be more honest, as long as their thresholds have actually been tested.

Multiple signals reduce specific failure modes

Semantic embeddings can over-connect queries that share a topic but not a task. Lexical features can miss paraphrases entirely. SERP overlap can be volatile week to week. Entity matching can be too strict if aliases have not been resolved. The benefit of a hybrid graph is not that all these errors disappear. It is that one signal can challenge another and catch mistakes a single signal would miss.

A practical edge-building pipeline

  1. Generate semantic candidate neighbours.
  2. Calculate lexical evidence for those candidates.
  3. Extract and compare entities.
  4. Attach SERP evidence where available.
  5. Apply hard incompatibility rules.
  6. Scale and combine the remaining signals.
  7. Keep the component values for explanation.
  8. Evaluate downstream clusters with and without each signal.
Practitioner principle: a strong weighted graph should be easier to explain than a semantic-only graph, not harder.

Worked example: when the signals disagree

Take "best project management software" and "project management software pricing". Semantic similarity might be very high, and the entity evidence might fully agree. But SERP overlap could still be modest, because comparison pages dominate one result set while vendor pricing pages dominate the other. That disagreement is genuinely useful information. Rather than averaging everything into a medium score and losing the reason behind it, preserve the conflict as a review state: same topic, different likely page job.

Calibrate weights against decisions, not aesthetics

A graph can look tidier after you increase the semantic weight, but that alone does not prove the change is an improvement. Tune weights against a judgement set built from the kinds of mistakes that actually matter to your site. For URL mapping, false merges may cost more than false separations. For topic discovery, broader recall might be preferable instead.

Keep signal provenance attached to every edge

If ClusterIQ outputs a final edge state such as strong, moderate or review required, the underlying signal values should stay attached to that relationship. That way, a future model change can be compared against the same semantic, lexical, entity and SERP evidence, rather than just the final combined weight. This matters most when practitioners question a surprising community: they need to see whether the edge exists because several signals agree, or because one heavily weighted component simply dominated the decision.

ClusterIQ Conclusion

Weighted keyword graphs are useful precisely because SEO relationships are multi-dimensional. The strongest design keeps the evidence components separate, uses hard constraints where a distinction genuinely matters, and calibrates weights against the actual task. That gives you a graph that benefits from several signals without turning into an opaque scoring system.

Sources and further reading

Farky Rafiq

Farky Rafiq

Founder of ClusterIQ

I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.

Put the idea into practice with your own keyword data

ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.