PCA vs UMAP for keyword embeddings: reduction for analysis, speed and visualisation
PCA and UMAP both reduce high-dimensional embeddings, but they preserve different structures. Learn when each method helps keyword analysis and when reduction changes the problem.

Farky Rafiq
Founder of ClusterIQ

High-dimensional embeddings are powerful, but they're awkward to work with day to day.
A sentence model might represent every keyword with hundreds of values. That's fine for a computer to crunch through, but hard for a person to visualise, and some clustering or neighbour-search workflows can actually benefit from a smaller representation.
PCA and UMAP are the two most common reduction methods for this. They are not interchangeable.
PCA is linear
Principal Component Analysis finds the directions that explain as much variance as possible, then projects the data onto a smaller set of orthogonal components along those directions.
Scikit-learn describes PCA as linear dimensionality reduction using singular value decomposition.
That makes PCA relatively easy to reason about. It's deterministic under common solver settings, and its components have a direct, traceable relationship to variance in the original data.
UMAP is neighbourhood-oriented
UMAP builds a lower-dimensional representation designed to preserve important aspects of the local structure in your data.
That makes it particularly good for visualisation, since dense neighbourhoods and broad topical regions can become visible once you plot it in two dimensions.
As we discuss in our UMAP guide, the picture you get is still a projection. Parameters like n_neighbors and min_dist change what the viewer actually sees.
PCA is often useful before expensive downstream work
If your original embedding has 768 dimensions, reducing it to 100 or 200 with PCA can cut storage and computation costs while keeping much of the variance intact.
That's useful ahead of:
- nearest-neighbour search;
- density clustering;
- exploratory analysis;
- model comparison.
The right number of dimensions to keep should be tested, not chosen simply because it's convenient.
UMAP can reveal non-linear local structure
PCA can only build linear combinations of the original dimensions. UMAP can capture more complex local relationships that a straight line can't.
That makes UMAP useful where the embedding space has curved or non-linear structure that a linear projection would flatten out badly.
The trade-off is more sensitivity to parameters and a stronger need for reproducibility controls.
Two dimensions are usually for people, not for clustering
A common workflow looks like this:
embeddings → UMAP 2D → HDBSCAN
It can produce attractive-looking results, but that two-dimensional projection was chosen because it's easy to plot, not necessarily because it preserves the best structure for clustering.
A stronger experiment compares:
- clustering in the original embedding space;
- clustering after a moderate PCA reduction;
- clustering after a higher-dimensional UMAP reduction;
- a separate two-dimensional UMAP kept purely for display.
Variance preserved is easier to inspect with PCA
PCA gives you explained variance ratios, a concrete way to check how much variance the retained components actually capture.
That's not the same as preserving semantic usefulness, but it's a transparent starting point for choosing how many dimensions to keep.
UMAP doesn't offer an equivalent simple "percentage of semantic information retained" figure.
Reproducibility differs between the two
UMAP relies on stochastic optimisation, and its documentation covers how random state affects reproducibility and performance.
PCA is generally much easier to reproduce under a fixed solver and software environment.
If UMAP is part of a production clustering pipeline, keep a record of:
- version;
- random state;
n_neighbors;min_dist;- metric;
- output dimensions.
Don't optimise the reduction independently of the task
A two-dimensional map that looks beautifully separated on screen may not actually produce the best page-level clusters.
Equally, retaining 95% of PCA variance doesn't prove the representation preserves the entity distinctions an ecommerce taxonomy needs.
Evaluate any reduction through downstream checks instead:
- nearest-neighbour quality;
- cluster stability;
- judgement-set accuracy;
- entity preservation;
- speed and memory cost.
A practical decision rule
Use PCA when you want a transparent, linear compression, and speed, storage or reproducibility in a preprocessing stage matter to you.
Use UMAP when local neighbourhood structure and visual exploration matter, or when your experiments show a non-linear representation actually improves the clustering task.
Use both when the jobs are genuinely different: PCA for production reduction, UMAP 2D for practitioner visualisation.
Practitioner principle: dimensionality reduction should simplify the representation, not quietly redefine which keywords count as related.
ClusterIQ Conclusion
PCA and UMAP are both genuinely useful tools for keyword embeddings, but they make different assumptions under the hood.
PCA gives you a linear, variance-oriented representation. UMAP emphasises neighbourhood structure and shines for visualisation. Test both as components of the full pipeline, and keep display-oriented projections separate from production clustering unless the evidence clearly supports combining them.
A two-stage reduction can be sensible
For very large embedding sets, PCA and UMAP don't have to compete. PCA can first reduce a 768-dimensional embedding down to a more manageable intermediate size, stripping out low-variance dimensions and cutting computation. UMAP can then build a lower-dimensional neighbourhood representation on top of that for exploratory analysis.
The key point is that every transformation becomes part of the model. If clustering happens after both steps, changing either the PCA dimensionality or the UMAP parameters can change the resulting groups.
Keep a no-reduction baseline
Whenever you introduce reduction for speed or visualisation, keep a baseline run in the original embedding space too. If the reduced pipeline produces substantially different nearest neighbours or cluster boundaries, dig into why. Reduction earns its place when it removes redundancy without removing the distinctions your SEO task actually needs.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

UMAP for SEO: visualising keyword embeddings without mistaking the map for the data

How to cluster 100,000 keywords without comparing every pair
