Building a Search Console query-page matrix for cluster analysis
Search Console can connect queries to the pages Google already shows. A query-page matrix turns that data into evidence for page ownership, cannibalisation and cluster validation.

Farky Rafiq
Founder of ClusterIQ

You have grouped 800 keywords into clusters, but which pages should you improve, combine or create? Before changing the site, it helps to see which URLs already appear in Google for those queries.
Keyword research tells you what people may search for. Search Console tells you which queries have actually shown pages from your property in Google Search.
A query-page matrix brings those queries and URLs together in a table. After clustering, it lets you check whether your groups of related keywords match the way your existing pages appear in search.
What the matrix represents
Each row is a query and each column is a URL. Where a row and column meet, the cell can contain:
- impressions;
- clicks;
- click-through rate, or CTR;
- average position;
- or a simple indicator that the URL appeared for that query.
Google's Performance report supports both query and page dimensions. Read the results carefully, though: the rules for combining data change depending on the dimension you select.
Join cluster IDs onto the query rows
Give each query its cluster ID, the label that identifies its group, and you can summarise the table at cluster level.
For each cluster, calculate:
- how many URLs appear for its queries;
- which URL receives the most impressions or clicks;
- what share of those impressions or clicks that leading URL receives;
- which other URLs appear;
- which queries have no consistent leading page.
This gives you a useful view of page ownership: which page, if any, is the main destination for a group of related queries.
A strong dominant URL is often reassuring
If one page receives most impressions across a coherent cluster of related queries, and its purpose matches that cluster, your current page mapping may already be working well.
The right next step could be to improve that page rather than create a new one.
Several strong URLs need interpretation
If two or more pages share visibility within a cluster, investigate why before deciding that something needs fixing.
Possible explanations include:
- different page types serving different search intents;
- valid parent-child relationships, such as a category and a subcategory;
- location or market differences;
- duplicate content left over from earlier site changes;
- unclear page ownership.
The matrix shows you the pattern. It cannot explain the cause on its own.
Use time as another dimension
Build a matrix for each month to see whether the same pages consistently lead their clusters.
Track:
- changes in the dominant URL;
- new URLs appearing within the cluster;
- older URLs losing visibility;
- signs of cluster demand rising or falling.
This is particularly useful when reviewing keyword cannibalisation. Two pages appearing alongside each other is different from pages repeatedly swapping places as the main result.
See our cannibalisation guide for the decision framework.
Search Console is incomplete by design
Google notes that some query data is anonymised and that filtering can affect totals. The Performance report also uses different rules when aggregating data by page or by property.
So do not treat your matrix as a complete record of every search or every URL that ranked.
It is first-party evidence from your property. That makes it useful, but not exhaustive.
Use cluster-level metrics carefully
Adding up query impressions within a cluster can help you prioritise work. Be clear about the date range and how you have aggregated the data.
A cluster containing many long-tail queries, often more specific searches, may look large because your dataset includes many related rows. Assess demand alongside:
- the number of unique queries;
- page ownership;
- clicks;
- commercial relevance;
- the spread of current ranking positions.
The matrix can validate new clusters
If a cluster looks coherent based on the meaning of its queries, and those queries already lead to one page, that provides useful supporting evidence for the grouping.
If the queries are spread across several clearly different page types, the cluster may be too broad to map to a single URL.
This is a useful way to test a clustering model's output against independent evidence, rather than accepting every group at face value.
It can also reveal missing pages
A group of related queries might generate impressions for several weak or only loosely relevant URLs, with no clear leading page.
That may point to a content gap. Before creating anything, check that the group represents a distinct user task and that a new page would serve people better than an improved existing page.
A practical implementation
- Export Search Console query-page data for a defined period.
- Normalise URLs conservatively, making their format consistent without merging genuinely different pages.
- Match query rows to their cluster IDs.
- Build summaries showing the URLs associated with each cluster.
- Measure the dominant URL's share and identify competing URLs.
- Compare the results across time periods.
- Flag unstable or high-value clusters for review.
- Record the eventual decision: keep, improve, consolidate or create.
Practitioner principle: semantic clustering tells you which queries look related. The query-page matrix shows how your site is already being interpreted for those queries.
ClusterIQ Conclusion
A Search Console query-page matrix is a useful bridge between keyword clustering and decisions about your site's pages and structure.
It helps you check your groups, spot unclear page ownership, monitor URL switching and prioritise possible content gaps using first-party evidence.
Keep Search Console's aggregation rules and privacy limitations in mind. Use the matrix to inform page decisions, not as a complete picture of everything happening in search.
Worked example: one cluster, three URLs
Suppose a cluster around “keyword clustering software” receives 70% of its impressions through a product page, 20% through a blog guide and 10% through an old landing page. That split is not automatically a problem.
If the blog leads for queries with informational wording and the product page leads for commercial queries, the split may be healthy. If all three URLs appear for the same commercial queries and take turns leading from month to month, ownership is much less clear.
Normalise time before comparing pages
Use consistent date windows when calculating the dominant URL's share. Comparing one URL's seasonal peak with another URL's quiet period can lead you to the wrong conclusion. For large sites, monthly cluster-page summaries give you a compact historical record of how page ownership changes.
Sources and further reading

Farky Rafiq
Founder of ClusterIQ
I've worked in digital marketing since 2005 and founded Liquid Silver in 2011. These articles are where I share the methods, experiments and practical SEO thinking behind ClusterIQ.
Put the idea into practice with your own keyword data
ClusterIQ helps turn raw SEO exports into clean, structured working datasets you can inspect, refine, report on and take into the next stage of your workflow.
Keep reading

Keyword cannibalisation: how clustering helps find overlap without inventing a problem

Mapping keyword clusters to existing URLs with embeddings
