BERTopic for SEO: useful topic discovery, but not an automatic content plan
BERTopic combines embeddings, clustering and c-TF-IDF to discover topics. That can be useful for SEO research, but the resulting topics still need editorial and page-level interpretation.

Farky Rafiq
Founder of ClusterIQ
Run BERTopic on a pile of queries and you get something that looks instantly useful: a list of topics, a set of representative words for each one, and example documents to check against. It is tempting to treat that output as a finished content plan. Resist that temptation.
BERTopic is a modular topic-modelling framework, not a single algorithm. Its default pipeline combines document embeddings, dimensionality reduction, clustering and class-based TF-IDF, and each stage brings its own assumptions into the final topic.
What BERTopic actually does
At a high level, the default workflow is:
- embed documents or queries;
- reduce the embedding space, commonly with UMAP;
- cluster the reduced points, commonly with HDBSCAN;
- represent each cluster using c-TF-IDF;
- extract representative words and documents.
This is useful because the semantic grouping stage and the human-readable labelling stage are kept separate. You can inspect one without the other breaking.
c-TF-IDF makes the clusters easier to name
BERTopic's class-based TF-IDF treats every document in a topic as one combined class, then finds the terms that distinguish that class from the others. For SEO work, this helps with:
- cluster labels;
- topic summaries;
- distinctive terminology;
- comparisons between neighbouring topics.
We rely on the same underlying idea in our cluster-labelling workflow.
A topic and a page are not the same thing
A single BERTopic topic can contain several different user intents or page formats at once. A topic labelled "project management software" might mix:
- pricing queries;
- comparison queries;
- implementation advice;
- templates;
- reviews.
That is a genuinely useful semantic topic. It just is not necessarily one page. This is the same distinction we cover in keyword clustering versus topic clustering: broad topic discovery and URL-level grouping sit at different levels of abstraction, and mixing them up leads to pages that try to answer too many questions at once.
The pipeline is configurable, which is a blessing and a trap
BERTopic lets you swap out the embedding model, the dimensionality-reduction model, the clustering model, the vectoriser, the topic representation and the outlier handling. That flexibility is powerful, but it also means "we used BERTopic" tells you almost nothing on its own. Record the full configuration you actually ran, or nobody will be able to reproduce your results later.
Topic size can mislead you
A large topic might reflect genuinely broad demand. It might also reflect a clustering stage that is simply too permissive. A small topic could be a specialist but commercially valuable concept rather than noise. Do not prioritise content purely because a topic looks big. Add demand data, business value, existing coverage and page-type evidence before deciding what to build.
Representative words can hide the entities that matter
Frequency-based topic terms tend to favour generic, recurring language. If a cluster contains several distinct product models, the top words might describe the category well while quietly omitting the exact entities that determine which page should own which query. Keep entity extraction and representative documents available alongside the topic words, not instead of them.
Where BERTopic earns its keep
It becomes especially useful once your input is broader than a tidy keyword list, for example:
- Search Console queries;
- support tickets;
- reviews;
- competitor headings;
- site-search data;
- article corpora.
The model can surface recurring semantic regions without needing a fixed list of categories decided in advance.
Treat topics as a discovery layer, not a final answer
A sensible SEO workflow looks like this:
- run topic discovery;
- inspect representative terms and documents;
- rename topics using language practitioners actually use;
- compare against existing categories and pages;
- split mixed-intent topics where needed;
- merge trivial or duplicated topics;
- decide which topics deserve content, taxonomy changes or no action at all.
The model speeds up discovery. The site architecture is still a human decision.
Do not rush to eliminate outliers
Density-based pipelines will leave some observations outside the main topics. BERTopic includes outlier-reduction techniques that reassign these points based on other evidence, and that can be genuinely useful. But forced reassignment should never become the goal in itself. Some outliers are novel, specialist or simply ambiguous, and keeping them visible for review tells you more than pretending everything fits somewhere.
Test the model against a human topic map
One of the strongest checks is comparing discovered topics against an independent practitioner map. Ask yourself:
- Which known topics did the model recover?
- Which model topics are genuinely new?
- Where did it merge concepts you would keep separate?
- Where did it fragment one useful topic into several?
- Which labels are misleading?
Used this way, the model challenges your thinking rather than replacing it.
Practitioner principle: BERTopic is most valuable when it shows you a structure worth investigating, not when its topic numbers get copied straight into a publishing calendar.
Worked example: one topic, several page jobs
Say BERTopic discovers a topic whose strongest terms are "email marketing", "automation", "software", "pricing" and "templates". The representative queries might mix software comparisons, pricing research and requests for downloadable templates. The model has correctly found a broader email-marketing topic. The mistake would be turning that single topic into a single article. Instead, keep the topic as the parent concept, then use intent and page-type evidence to split out comparison, pricing and template needs underneath it.
Use topic reduction carefully
BERTopic can merge similar topics to cut down fragmentation, which is helpful for exploration. But topic reduction can also erase commercially important distinctions. Check which topics are being merged and whether their representative documents still support one coherent editorial purpose before you accept the merge.
ClusterIQ Conclusion
BERTopic combines several useful techniques into a practical topic-discovery framework. Its embeddings, density clustering and c-TF-IDF representations can surface structure in large text collections quickly. For SEO, that output belongs upstream of content planning, not in place of it. Validate the topics against intent, page type, entities, existing coverage and business priorities before turning any of them into URLs or articles.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Adding new keywords to existing clusters without rebuilding everything

HDBSCAN for keyword clustering: density, noise and uncertain queries
