TF-IDF vs semantic embeddings for keyword clustering
TF-IDF and semantic embeddings preserve different signals. Learn where each representation works, where it fails and why hybrid keyword clustering can be stronger.

Farky Rafiq
Founder of ClusterIQ

Two keywords can look almost identical and still need different pages. Two others can use completely different words and express the same need.
Suppose your 500-keyword Semrush export includes “project management software pricing” and “project management template download”. The wording overlaps, but one person wants to compare a product and the other wants a resource. Elsewhere, “cheap running trainers” and “affordable running shoes” look different but probably belong in the same conversation.
This is the practical problem behind TF-IDF and semantic embeddings. Both help a clustering system decide which queries are related. They pay attention to different evidence, so they also make different mistakes. The right question is not which method sounds more modern. It is which one preserves the distinctions your SEO plan needs.
TF-IDF pays close attention to the words
TF-IDF stands for term frequency-inverse document frequency. The name is technical, but the idea is useful: words that are distinctive within your dataset should carry more weight than words that appear almost everywhere.
If a list contains hundreds of headphone queries, a common word such as “headphones” tells you relatively little. A combination such as “wireless”, “noise cancelling” and “children” may be much more informative. TF-IDF turns each query into a list of weighted word features so those patterns can be compared.
Scikit-learn's TfidfVectorizer can work with whole words, character sequences or n-grams, which are short runs of consecutive words or characters. With L2 normalisation, cosine similarity between TF-IDF vectors is equivalent to their dot product. In everyday use, cosine similarity is simply a way to measure how closely two weighted profiles point in the same direction.
The resulting representation is sparse because most queries contain only a small fraction of the full vocabulary. It is also inspectable. If two queries are close because they share “wireless”, “noise cancelling” and “headphones”, an SEO can see the reason.
Why exact wording still matters
It is easy to treat word-based methods as old-fashioned now that semantic models are widely available. In keyword research, that would be a mistake.
Consider these phrases:
- “iPhone 17 case”;
- “iPhone 17 Pro case”;
- “iPhone 17 Pro Max case”.
They are very close in meaning, but the model names are commercially important. A retailer may need separate categories or product filters. A method that smooths over “Pro” and “Pro Max” could create a visually neat cluster that leads to the wrong page plan.
TF-IDF often performs well when rare tokens, product codes, dimensions, materials, brands and technical modifiers matter. Character n-grams can also help with spelling variations and related word forms. This makes lexical, or word-based, evidence valuable in ecommerce and specialist B2B research.
Where word matching misses the point
The weakness appears when people describe the same task with different language:
- “how to reduce page load time”;
- “ways to make a website faster”.
A basic word-based comparison sees limited overlap. An experienced marketer sees the shared need immediately.
Stemming, lemmatisation and n-grams can help. Stemming and lemmatisation reduce related word forms towards a common base, while n-grams preserve short combinations. These are useful improvements, but they still work mainly with the words present. They do not provide broader knowledge of meaning in the way a trained language model can.
Embeddings look for meaning beyond exact wording
A semantic embedding represents a query as a dense numerical vector. “Dense” means most of its dimensions contain a value, unlike the mostly empty vocabulary profile produced by TF-IDF. Queries used in similar contexts and expressing similar ideas tend to be placed nearer one another.
Sentence-BERT was designed to create reusable sentence representations that can be compared efficiently. That made semantic search and clustering far more practical than comparing every sentence pair through the original BERT architecture.
For keyword work, embeddings can connect paraphrases such as:
- “cheap running trainers”;
- “affordable running shoes”.
This is their main attraction. A list from Search Console or Ahrefs contains the varied language real people use. Embeddings can recover relationships that exact word overlap would miss.
Meaning alone does not decide the page
Embeddings have the opposite failure mode. They may connect queries because the general subject is similar even when the required outcome differs.
Compare:
- “project management software pricing”;
- “how to manage agency projects”;
- “project management template download”.
All three belong to the wider subject of project management. They may still require a pricing page, an educational guide and a downloadable template. If they are forced into one cluster, the content brief becomes muddled and the page recommendation becomes difficult to use.
Semantic similarity therefore cannot replace intent and page-format analysis. Our article on keyword clustering versus topic clustering explores the wider distinction: being related is not the same as being suitable for one operational group.
The trade-off is precision, recall and explanation
TF-IDF tends to protect exact terminology. That can improve precision when a small modifier changes the product or task, but it may fail to recall paraphrases that should have been connected.
Embeddings tend to improve recall across varied language. They can also connect broad neighbours that an SEO would separate by intent, audience or page type.
There is also a difference in explainability. TF-IDF features are readable words and n-grams. Dense embedding dimensions do not normally map to labels a person can inspect. You can review nearest neighbours and examples, but you cannot usually say that dimension 214 means “commercial intent”.
That does not make embeddings unreliable. It means a good workflow should show supporting evidence:
- examples of each query's nearest neighbours;
- the spread of similarity scores;
- confidence or stability measures;
- the model and settings used;
- boundary cases that need human review.
That evidence matters when clustering will affect page ownership, content briefs or information architecture. A manager or client should be able to see why a surprising relationship appeared.
How the methods behave as a dataset grows
Both approaches can work efficiently on the few hundred or few thousand keywords found in a normal campaign, but they store information differently.
TF-IDF creates sparse vectors. Most query and term combinations are zero, which can be efficient even with a large vocabulary. Embeddings create dense vectors with a fixed number of dimensions set by the model.
The Sentence-BERT research addressed an important practical problem. Comparing sentence pairs directly with pre-trained BERT was expensive because every pair had to be processed together. SBERT creates each sentence embedding once, after which the vectors can be compared far more efficiently.
At much larger scales, representation is only part of the cost. Comparing every query with every other query grows quadratically. Approximate nearest-neighbour search, sparse graphs or candidate-generation steps can narrow the comparisons. For a 700-keyword plan, however, the more important concern is usually whether the relationships support good SEO decisions.
Using both often gives a better working answer
You do not have to declare a winner. A hybrid workflow can keep lexical and semantic evidence together:
- clean obvious formatting issues and combine exact duplicates;
- calculate TF-IDF or n-gram features;
- calculate semantic embeddings;
- find plausible semantic neighbours;
- strengthen or reject relationships using important words, entities and modifiers;
- send disagreements to review or a second classification step.
This protects against two expensive errors: separating queries that express the same need, and combining queries that share a broad topic but require different pages.
Graph methods can be useful because a relationship can carry more than one signal. The network does not have to pretend semantic similarity is the only evidence. Our guide to graph theory and keyword clustering explains that approach.
Which signal deserves more weight?
- Lean more on TF-IDF when model numbers, product attributes, brand names, locations or technical wording are decisive.
- Lean more on embeddings when the dataset contains many paraphrases, natural-language questions and varied descriptions of the same task.
- Use both when broad semantic variation sits alongside commercially important detail, which is common in real keyword exports.
For an informational content plan, embeddings may reveal topics hidden by different wording. For ecommerce or a specialist product taxonomy, lexical signals can stop critical attributes disappearing into a broad semantic average.
Test the method on the decisions you need to make
Do not judge a representation from three impressive examples. Create a small set of known cases from your own market. Include easy matches, paraphrases, near misses, product or entity changes and shifts in intent.
Then ask:
- Which genuine relationships does each method recover?
- Which false relationships does it create?
- Does it preserve modifiers that change the product or page?
- Are the resulting groups stable?
- Can a practitioner understand a borderline decision?
- Does the result make the page plan and content briefs clearer?
The best representation is the one that preserves the distinctions your work depends on, not the one with the newest architecture.
ClusterIQ Conclusion
TF-IDF and semantic embeddings solve different parts of the same practical problem. TF-IDF keeps the wording visible and protects exact details. Embeddings recognise meaning across synonyms and paraphrases.
Neither one understands search intent, page format or commercial priority by itself. Those decisions still need evidence and practitioner judgement.
For many keyword projects, a hybrid approach is the most useful. Semantic evidence broadens what the system can see, while lexical evidence protects the details that determine whether two queries really belong on the same page.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Keyword preprocessing before clustering: clean the data without erasing intent
Entity-aware keyword clustering: preserving brands, products and attributes
