WritingCopywritingAdvanced16 minSaves 180 minutes

A/B Copy Test Hypothesis Builder

Turn copy ideas into falsifiable hypotheses with honest sample size estimates.

Turns copy change ideas into falsifiable test hypotheses with the belief each challenges, the variant to build, metrics, a realistic sample estimate, and a plain statement of which tests your traffic cannot support.

Ready-to-use prompt

The prompt

Copy it as-is, then swap the bracketed placeholders for your own details before running it.

prompt.txt
Role: You are a conversion optimisation strategist writing test hypotheses.

Context:
- Page or asset and its current copy: {{asset}}
- Current conversion rate and monthly volume: {{performance}}
- Qualitative signals, session recordings, surveys, support themes: {{signals}}
- Changes being considered: {{ideas}}
- Testing tool and constraints: {{tooling}}

Task: For each idea return a hypothesis in this form:
1. Because we observed [evidence], we believe [change] will cause [effect] for [segment], measured by [metric]
2. The underlying belief about the audience that this test challenges
3. The exact variant copy to build, written out
4. Primary metric and one guardrail metric
5. Minimum sample per variant at the current baseline, with the assumed effect size stated
6. Estimated runtime at the supplied traffic volume
7. What you would do if it wins, loses, or comes out flat

Then return:
- Hypotheses ranked by expected learning value, not by expected lift
- Any hypothesis this traffic volume cannot reach significance on, said plainly
- Ideas that should be shipped without a test because the downside is negligible
- Two tests that must not run at the same time, and the interference between them

Rules:
- Every hypothesis must be falsifiable, reject anything that cannot lose.
- Never propose a test that changes several variables at once unless it is explicitly a bundle test, and label it.
- Be honest about underpowered tests instead of estimating them optimistically.

- Separate supplied evidence, assumptions and missing inputs; never fabricate research observations.
- State allocation, significance level, power, calculation method, and absolute versus relative effect size.
- Base runtime on eligible unique units; include outcome maturation and the tool’s stopping rule.
- Distinguish evidence of negligible effect from inconclusive results using uncertainty intervals.

Finish with: the single test to run first, and what you learn either way.

Estimated results

DifficultyAdvanced
Setup time16 min
Time saved180 minutes
Best modelsChatGPT, Gemini, Claude
Best audienceSaaS, Ecommerce

Editor's note

Why this prompt matters

Most testing roadmaps are lists of ideas with the word test in front of them. They have no stated belief, no falsifiable outcome and no honest view of whether the traffic can decide the question, which is why so many programmes produce a year of inconclusive results and quiet abandonment. This prompt turns each idea into a hypothesis that can lose, attaches the metric that decides it, and — the part no testing tool volunteers — says plainly which experiments this site will never be able to call.

Anatomy

Prompt engineering breakdown

Role

Context

Goal

Constraints

Output format

What you'll get

Expected output

Worked example: trial signup page

This fictional planning example uses explicit assumptions, not measured results: 12,000 eligible unique visitors monthly, a 4% visitor-to-signup rate, equal allocation, and an eight-week testing window. The product genuinely requires no credit card. Illustrative research themes are payment anxiety and uncertainty about setup effort. The current button says “Start free trial.”

Hypotheses ranked by learning value

1. Remove payment uncertainty. Because the illustrative research mentions unexpected charges, we believe adding payment reassurance will increase completed signups among new visitors, measured by visitor-to-signup conversion.

  • Belief challenged: Visitors already understand that starting a trial creates no payment commitment.
  • Exact variant: Keep “Start free trial” unchanged; add “No credit card required.” immediately below it. Change nothing else.
  • Primary metric: Completed signups per assigned visitor.
  • Guardrail: Signup form error rate; agree an unacceptable increase before launch.
  • Decision: If it wins without guardrail harm, retain the reassurance. If it loses, remove it and investigate whether mentioning cards introduced anxiety. If inconclusive, retain control; a wide interval does not establish that payment concerns are irrelevant.

2. Make the next step concrete. Because the illustrative research suggests uncertainty about setup, we believe replacing the button label with “Set up my workspace” will increase completed signups among new visitors, measured by visitor-to-signup conversion.

  • Belief challenged: Trial language motivates action better than a concrete next step.
  • Exact variant: Replace only “Start free trial” with “Set up my workspace”.
  • Primary metric: Completed signups per assigned visitor.
  • Guardrail: Workspace activation per assigned visitor within seven days.
  • Decision: If it wins, adopt after guardrail maturation. If it loses, restore control. If inconclusive, do not claim visitors prefer either framing.

Feasibility before scheduling

Approximate fixed-horizon, two-sided proportions calculations assume 5% significance and 80% power. Effect sizes are planning thresholds, not predictions.

| Test | Assumed detectable change | Visitors per arm | Recruitment runtime | |---|---|---:|---:| | Payment reassurance | 4% → 4.8%; 20% relative | 10,300 | About 7.5 weeks | | Concrete next step | 4% → 4.4%; 10% relative | 39,500 | About 29 weeks |

Only the first fits the window. Neither guarantees significance. Verify estimates against the testing tool’s method.

Recommendation

Run payment reassurance first: it probes whether payment uncertainty materially suppresses signup. Do not run both tests on overlapping visitors; both alter interpretation of the same action. Ship spelling corrections separately before launch. Defer the second test rather than promise an answer this traffic cannot support.

Under the hood

Why this prompt works

Ranking by learning value rather than expected lift is what makes a roadmap compound. A small win you cannot explain does not inform the next test; a clear loss on a stated belief does. Requiring falsifiability removes the ideas that cannot fail, which are always the most popular ones. And naming the underpowered tests before you run them is the difference between a programme that concentrates its traffic on questions it can answer and one that spreads it thinly across questions it cannot.

Model fit

Best AI models for this prompt

Claude

Best here. Most honest about underpowered tests and quickest to reject a non-falsifiable idea.

ChatGPT

Good at sample size arithmetic and at writing out the variant copy in full.

Gemini

Useful when you paste analytics exports alongside the ideas list.

When to use

  • After research synthesis, before drafting variants.
  • During roadmap planning when ideas exceed traffic capacity.
  • Before experiment approval, to expose unsupported sample assumptions.
  • When prioritising competing copy changes on one conversion path.

When not to use

  • Traffic cannot support a decision within your window.
  • Assignment or conversion tracking is unvalidated.
  • A redesign changes layout, offer and copy together.
  • Analytics baselines need verification by an analyst.
  • Regulated claims need qualified human review.

Get more from it

Pro tips

  • 1

    Rank learning value by naming the audience belief each result would update.

  • 2

    Keep the underpowered list; request required traffic before accepting a shorter runtime.

  • 3

    Ship spelling fixes without testing, but exclude pricing, consent and substantive claims from “negligible downside”.

  • 4

    Prewrite win, lose and inconclusive actions, including guardrail failures.

  • 5

    Supply eligible unique visitors, not pageviews; deduct excluded segments and competing experiments.

  • 6

    For a second pass, request recalculation at smaller detectable effects and comparison against your scheduling window.

Don't ship this

Common mistakes

  • Hypotheses that cannot lose.

    Fix — Require falsifiability and reject anything phrased as an improvement.

  • Running tests the traffic cannot decide.

    Fix — Use the sample estimate and the underpowered list.

  • Running two interfering tests together.

    Fix — Check the interference pairs before scheduling.

People also ask

Frequently asked questions

Q.What makes a hypothesis falsifiable?

It states an expected effect that can fail to appear. This will improve the page cannot lose, so it teaches nothing; a named metric moving for a named segment can be wrong, which is what makes it worth running.

Q.Why rank by learning value rather than expected lift?

Because a small unexplained win does not inform the next test, while a clear result on a stated belief does. Programmes that rank by lift plateau; programmes that rank by learning compound.

Q.What should I do with the underpowered list?

Ship those changes without testing if the downside is small, or combine them into a bigger swing. Running an experiment your traffic cannot decide spends weeks to produce a result you cannot act on.

Version 1.1Last reviewed September 21, 2026
Reviewed by editorial