A/B Copy Test Hypothesis Builder
Turn copy ideas into falsifiable hypotheses with honest sample size estimates.
Turns copy change ideas into falsifiable test hypotheses with the belief each challenges, the variant to build, metrics, a realistic sample estimate, and a plain statement of which tests your traffic cannot support.
Ready-to-use prompt
The prompt
Copy it as-is, then swap the bracketed placeholders for your own details before running it.
Role: You are a conversion optimisation strategist writing test hypotheses.
Context:
- Page or asset and its current copy: {{asset}}
- Current conversion rate and monthly volume: {{performance}}
- Qualitative signals, session recordings, surveys, support themes: {{signals}}
- Changes being considered: {{ideas}}
- Testing tool and constraints: {{tooling}}
Task: For each idea return a hypothesis in this form:
1. Because we observed [evidence], we believe [change] will cause [effect] for [segment], measured by [metric]
2. The underlying belief about the audience that this test challenges
3. The exact variant copy to build, written out
4. Primary metric and one guardrail metric
5. Minimum sample per variant at the current baseline, with the assumed effect size stated
6. Estimated runtime at the supplied traffic volume
7. What you would do if it wins, loses, or comes out flat
Then return:
- Hypotheses ranked by expected learning value, not by expected lift
- Any hypothesis this traffic volume cannot reach significance on, said plainly
- Ideas that should be shipped without a test because the downside is negligible
- Two tests that must not run at the same time, and the interference between them
Rules:
- Every hypothesis must be falsifiable, reject anything that cannot lose.
- Never propose a test that changes several variables at once unless it is explicitly a bundle test, and label it.
- Be honest about underpowered tests instead of estimating them optimistically.
- Separate supplied evidence, assumptions and missing inputs; never fabricate research observations.
- State allocation, significance level, power, calculation method, and absolute versus relative effect size.
- Base runtime on eligible unique units; include outcome maturation and the tool’s stopping rule.
- Distinguish evidence of negligible effect from inconclusive results using uncertainty intervals.
Finish with: the single test to run first, and what you learn either way.Estimated results
Editor's note
Why this prompt matters
Most testing roadmaps are lists of ideas with the word test in front of them. They have no stated belief, no falsifiable outcome and no honest view of whether the traffic can decide the question, which is why so many programmes produce a year of inconclusive results and quiet abandonment. This prompt turns each idea into a hypothesis that can lose, attaches the metric that decides it, and — the part no testing tool volunteers — says plainly which experiments this site will never be able to call.
Anatomy
Prompt engineering breakdown
Role
Context
Goal
Constraints
Output format
What you'll get
Expected output
Worked example: trial signup page
This fictional planning example uses explicit assumptions, not measured results: 12,000 eligible unique visitors monthly, a 4% visitor-to-signup rate, equal allocation, and an eight-week testing window. The product genuinely requires no credit card. Illustrative research themes are payment anxiety and uncertainty about setup effort. The current button says “Start free trial.”
Hypotheses ranked by learning value
1. Remove payment uncertainty. Because the illustrative research mentions unexpected charges, we believe adding payment reassurance will increase completed signups among new visitors, measured by visitor-to-signup conversion.
- Belief challenged: Visitors already understand that starting a trial creates no payment commitment.
- Exact variant: Keep “Start free trial” unchanged; add “No credit card required.” immediately below it. Change nothing else.
- Primary metric: Completed signups per assigned visitor.
- Guardrail: Signup form error rate; agree an unacceptable increase before launch.
- Decision: If it wins without guardrail harm, retain the reassurance. If it loses, remove it and investigate whether mentioning cards introduced anxiety. If inconclusive, retain control; a wide interval does not establish that payment concerns are irrelevant.
2. Make the next step concrete. Because the illustrative research suggests uncertainty about setup, we believe replacing the button label with “Set up my workspace” will increase completed signups among new visitors, measured by visitor-to-signup conversion.
- Belief challenged: Trial language motivates action better than a concrete next step.
- Exact variant: Replace only “Start free trial” with “Set up my workspace”.
- Primary metric: Completed signups per assigned visitor.
- Guardrail: Workspace activation per assigned visitor within seven days.
- Decision: If it wins, adopt after guardrail maturation. If it loses, restore control. If inconclusive, do not claim visitors prefer either framing.
Feasibility before scheduling
Approximate fixed-horizon, two-sided proportions calculations assume 5% significance and 80% power. Effect sizes are planning thresholds, not predictions.
| Test | Assumed detectable change | Visitors per arm | Recruitment runtime | |---|---|---:|---:| | Payment reassurance | 4% → 4.8%; 20% relative | 10,300 | About 7.5 weeks | | Concrete next step | 4% → 4.4%; 10% relative | 39,500 | About 29 weeks |
Only the first fits the window. Neither guarantees significance. Verify estimates against the testing tool’s method.
Recommendation
Run payment reassurance first: it probes whether payment uncertainty materially suppresses signup. Do not run both tests on overlapping visitors; both alter interpretation of the same action. Ship spelling corrections separately before launch. Defer the second test rather than promise an answer this traffic cannot support.
Under the hood
Why this prompt works
Ranking by learning value rather than expected lift is what makes a roadmap compound. A small win you cannot explain does not inform the next test; a clear loss on a stated belief does. Requiring falsifiability removes the ideas that cannot fail, which are always the most popular ones. And naming the underpowered tests before you run them is the difference between a programme that concentrates its traffic on questions it can answer and one that spreads it thinly across questions it cannot.
Model fit
Best AI models for this prompt
Claude
Best here. Most honest about underpowered tests and quickest to reject a non-falsifiable idea.
ChatGPT
Good at sample size arithmetic and at writing out the variant copy in full.
Gemini
Useful when you paste analytics exports alongside the ideas list.
When to use
- After research synthesis, before drafting variants.
- During roadmap planning when ideas exceed traffic capacity.
- Before experiment approval, to expose unsupported sample assumptions.
- When prioritising competing copy changes on one conversion path.
When not to use
- Traffic cannot support a decision within your window.
- Assignment or conversion tracking is unvalidated.
- A redesign changes layout, offer and copy together.
- Analytics baselines need verification by an analyst.
- Regulated claims need qualified human review.
Get more from it
Pro tips
- 1
Rank learning value by naming the audience belief each result would update.
- 2
Keep the underpowered list; request required traffic before accepting a shorter runtime.
- 3
Ship spelling fixes without testing, but exclude pricing, consent and substantive claims from “negligible downside”.
- 4
Prewrite win, lose and inconclusive actions, including guardrail failures.
- 5
Supply eligible unique visitors, not pageviews; deduct excluded segments and competing experiments.
- 6
For a second pass, request recalculation at smaller detectable effects and comparison against your scheduling window.
Don't ship this
Common mistakes
✗ Hypotheses that cannot lose.
Fix — Require falsifiability and reject anything phrased as an improvement.
✗ Running tests the traffic cannot decide.
Fix — Use the sample estimate and the underpowered list.
✗ Running two interfering tests together.
Fix — Check the interference pairs before scheduling.
People also ask
Frequently asked questions
Q.What makes a hypothesis falsifiable?
It states an expected effect that can fail to appear. This will improve the page cannot lose, so it teaches nothing; a named metric moving for a named segment can be wrong, which is what makes it worth running.
Q.Why rank by learning value rather than expected lift?
Because a small unexplained win does not inform the next test, while a clear result on a stated belief does. Programmes that rank by lift plateau; programmes that rank by learning compound.
Q.What should I do with the underpowered list?
Ship those changes without testing if the downside is small, or combine them into a bigger swing. Running an experiment your traffic cannot decide spends weeks to produce a result you cannot act on.