Keyword Clustering Into Page-Level Topic Groups
Collapse a messy keyword export into clusters that each map to exactly one page.
Takes a raw keyword list and clusters it by search intent and SERP overlap into groups that each justify a single page, naming a primary keyword, secondary keywords and the page type for every cluster.
Ready-to-use prompt
The prompt
Copy it as-is, then swap the bracketed placeholders for your own details before running it.
Role: You are an SEO strategist clustering keywords into page-level topics.
Context:
- Keyword list: {{keyword_list}}
- Site type: {{site_type}}
- Existing site structure: {{site_structure}}
Task: Cluster the keywords so that every cluster maps to exactly one page. For each cluster return:
1. Cluster name
2. Primary keyword (the one the page targets in the title)
3. Secondary keywords (all remaining members)
4. Shared search intent
5. Page type (pillar, sub-topic, comparison, tool, product, glossary)
6. Suggested URL slug
7. Reason these keywords belong on one page rather than several
Rules:
- Cluster by what the searcher wants, not by shared words. "cheap X" and "X pricing" can be one cluster; "X reviews" and "X login" cannot.
- Every keyword must appear in exactly one cluster. List anything you cannot place under UNASSIGNED with a reason.
- Split a cluster whenever two keywords would need materially different page content.
- Cap each cluster at one primary keyword.
- Label clusters intent_inferred unless supplied ranking data supports SERP-overlap claims.
- Preserve keyword wording; report exact duplicates removed before clustering.
- Identify existing destination URLs before proposing new ones.
- Apply a declared, consistent tag vocabulary to every cluster member.
Finish with: a table of clusters ordered by how much of the list each one absorbs.Estimated results
Editor's note
Why this prompt matters
A keyword export is not a content plan. The gap between the two is clustering, and it is the step most teams do badly — grouping by shared words instead of shared intent, which is how you end up with four thin pages fighting each other for the same result. This prompt forces the harder question for every group: would one page genuinely satisfy every keyword in it? Anything that fails that test gets split, anything that cannot be placed goes to an explicit unassigned bucket instead of being quietly dropped.
Anatomy
Prompt engineering breakdown
Role
Context
Goal
Constraints
Output format
What you'll get
Expected output
Worked example: employee scheduling software
A generic scheduling-software site supplies 17 keywords. Its existing structure includes a product page at /employee-scheduling/ and a pricing page at /pricing/, but no template resource or glossary. The export contains no ranking URLs, so these are intent-based draft clusters, not verified SERP-overlap groups.
The result is four page-level clusters covering 16 keywords, plus one unresolved query. Two clusters map to existing pages; they are not recommendations to publish competing URLs. The primary keywords describe each page’s central purpose, rather than an assumed search-volume winner.
Decisions that affect the page plan
The product and pricing clusters stay separate because evaluating scheduling features and checking subscription costs require different main content. Pricing terms belong together here because one pricing page can explain plans, billing and total cost.
The template cluster assumes the proposed resource includes an editable spreadsheet with a weekly layout. Without that deliverable, the spreadsheet query would need reassessment rather than being forced into a generic article. Regional wording such as “rota” does not justify another page unless the target market or results reveal a different need.
Tagging and labelling inside a cluster
Use four fixed fields: intent, funnel_stage, persona and page_type. For this export, allow commercial, transactional or informational for intent; awareness, consideration or decision for funnel stage; and operations_manager or unknown for persona. Page-type tags use the prompt’s categories, normalised to snake_case, including sub_topic.
Apply tags to every member, not just the primary keyword. All product-cluster members inherit commercial, consideration, operations_manager, product. Pricing members receive commercial, decision, operations_manager, sub_topic; template members receive transactional, consideration, operations_manager, tool; glossary members receive informational, awareness, unknown, glossary.
These are planning labels, not observed audience facts. A conflicting intent or page-type tag triggers a cluster review; a different persona label alone does not require another page.
UNASSIGNED
employee scheduling login — navigational intent, but no destination is identified. Confirm the intended service before routing it to an existing sign-in page; do not create an acquisition page.
Clusters ordered by keywords absorbed
| Cluster / count | Primary keyword | Secondary keywords | Shared intent | Page type | URL slug | Why one page | |---|---|---|---|---|---|---| | Software / 5 | employee scheduling software | staff scheduling software; employee scheduling app; staff rota software; online employee scheduling | Evaluate software | product | employee-scheduling | Same feature and workflow evaluation | | Pricing / 4 | employee scheduling software pricing | employee scheduling software cost; staff scheduling software price; employee scheduling subscription cost | Check costs | sub-topic | pricing | One plan-and-billing explanation answers all four | | Templates / 4 | employee schedule template | staff rota template; weekly employee schedule template; employee schedule spreadsheet | Obtain an editable schedule | tool | employee-schedule-template | One weekly spreadsheet supplies the requested resource | | Definition / 3 | what is employee scheduling | employee scheduling definition; staff scheduling meaning | Understand the concept | glossary | glossary/employee-scheduling | One definition and example resolve all three |
Under the hood
Why this prompt works
The one-page test is the mechanism. By requiring a justification for why keywords belong on the same page, the prompt makes over-merging visible instead of invisible, and the rule that every keyword lands in exactly one cluster prevents the silent deduplication that makes AI clustering look tidier than it is. Naming a single primary keyword per cluster gives each future page one title to optimise, which is what stops the cluster map turning back into cannibalisation later.
Model fit
Best AI models for this prompt
Claude
Best in class here. It reliably refuses to merge keywords with different intent, which is the exact failure mode of this task.
ChatGPT
Handles the largest lists comfortably. Chunk anything over 300 keywords into batches.
Gemini
Strongest when you also paste the URLs currently ranking for each keyword so it can cluster on observed SERP overlap.
When to use
- After cleaning an export, before estimating page count.
- Before consolidating thin pages targeting near-identical needs.
- While separating a pillar’s scope from supporting pages.
- Before briefing writers who need page boundaries and consistent keyword tags.
When not to use
- For fewer than 15 keywords; manual review is usually quicker.
- When live SERP overlap is mandatory but ranking data is unavailable.
- For paid-search ad groups governed by bidding and match types.
- When specialist terminology requires an expert to distinguish genuinely different needs.
Get more from it
Pro tips
- 1
Paste one keyword per line; keep SERP evidence separately keyed to each query.
- 2
Split mixed-page-type clusters unless one concrete deliverable satisfies every member.
- 3
Reprocess UNASSIGNED with missing context, not instructions to force placement.
- 4
Test large clusters against a page outline; split members needing different main content.
- 5
Supply existing URLs and their purposes to distinguish updates from new pages.
- 6
Audit coverage: clustered members plus UNASSIGNED must equal the deduplicated input.
- 7
For disputed merges, supply ranking URLs from the same market, device and collection window.
Don't ship this
Common mistakes
✗ Clustering by shared words rather than shared intent.
Fix — Require the justification column and reject clusters whose reason is only lexical overlap.
✗ Letting one cluster absorb half the list.
Fix — Split it by page type and re-run on the fragment.
✗ Ignoring the unassigned keywords.
Fix — Run a second pass — they are usually a real cluster nobody planned for.
People also ask
Frequently asked questions
Q.How is this different from a keyword grouping tool?
Tools cluster on lexical or SERP overlap. This prompt clusters on whether a single page could genuinely satisfy every keyword in the group, and it makes the model justify each grouping, so you can audit and override the reasoning.
Q.How large a keyword list can I paste in?
Up to roughly 300 keywords in one pass with most models. Beyond that, split the list, cluster each batch, then run a final merge pass over the cluster names to catch duplicates across batches.
Q.What should I do with the unassigned keywords?
Run them back through the prompt on their own. Leftovers are often a coherent cluster that did not fit the dominant theme, and that second pass regularly surfaces a page nobody had planned.