Experiment Title: Paid Social Incrementality: Geo-Holdout Test for Q3 Lead Generation
Objective: To quantitatively measure the incremental lift of paid social campaigns on purchase conversions and associated revenue within target markets, isolating true campaign contribution from organic baseline trends and other marketing efforts.
Hypothesis:
- Null Hypothesis (H0): There is no statistically significant difference in
purchase conversion rates or revenue per user between geo-holdout control groups (no paid social exposure) and test groups (paid social exposure). - Alternative Hypothesis (H1): Paid social campaigns will demonstrate a statistically significant positive incremental lift in
purchase conversion rates and revenue per user in test groups compared to control groups.
Experiment Design (detailed Geo-Holdout methodology): This experiment will employ a randomized geo-holdout design. Geographies (DMAs or zip code clusters) will be assigned to either a test group (exposed to paid social campaigns) or a control group (suppressed from paid social campaigns). Media suppression in control geos will be executed by excluding these locations from all paid social targeting settings. The experiment will run for a predetermined duration, after which key performance indicators will be compared between groups to ascertain incremental lift.
Geo Selection Methodology & Pairing Strategy: Geos will be selected based on historical stability in purchase conversion rates, comparable population densities, and similar past purchase volumes. We will aim for 10-15 geo pairs. Geos will be paired based on pre-period purchase conversion rates, average revenue per user, and population size, ensuring the closest possible match across these covariates. Pairing will be done using a Mahalanobis distance matching algorithm to minimize variance between pairs. Media market saturation will be considered to ensure effective suppression is feasible.
Sample Size Calculation Parameters & Justification:
- Minimum Detectable Effect (MDE): 2.5% increase in
purchase conversion rate. Justification: This MDE represents a financially meaningful lift given our average monthly ad spend of $500,000 and target CAC, making the test cost-efficient and actionable. - Statistical Power: 0.80. Justification: A power of 80% is standard practice, indicating an 80% chance of detecting a true effect if one exists.
- Significance Level (Alpha): 0.05. Justification: A 5% alpha level is standard, minimizing the risk of false positives (Type I error).
Test Duration Justification: Based on the MDE, power, and alpha, and historical weekly purchase volume variability, a test duration of 6 weeks is estimated to achieve sufficient sample size within each geo to detect the MDE. This duration also accounts for typical purchase conversion cycles and allows for potential lagged effects.
Primary Metrics (with definitions and calculation methods):
- Incremental `Purchase` Conversion Rate: (Test Group
Purchase Conversions / Test Group Unique Users) - (Control Group Purchase Conversions / Control Group Unique Users). - Incremental Revenue per User: (Test Group Total Revenue / Test Group Unique Users) - (Control Group Total Revenue / Control Group Unique Users).
Guardrail Metrics (with definitions):
- Organic Search Volume (by geo): Number of organic searches for brand and product keywords originating from test vs. control geos.
- Direct Traffic (by geo): Number of users navigating directly to the website from test vs. control geos.
- Customer Acquisition Cost (CAC) on other channels (by geo): The average cost to acquire a new
purchase customer through non-paid-social channels within test vs. control geos.
Data Collection & Reporting Cadence: Data will be collected daily from our analytics platform and ad platforms. Weekly reports will track primary and guardrail metrics, with a final comprehensive report and decision brief issued one week after the experiment concludes.
Decision Rule (including statistical test and threshold): The campaign will be deemed incrementally successful if the incremental purchase conversion rate and/or incremental revenue per user show a statistically significant positive difference (p < 0.05) between test and control groups, as determined by a paired t-test on the geo-level differences. If a significant positive lift is observed, the campaign will be scaled. If no significant lift or a negative lift is observed, the campaign will be paused or significantly re-evaluated for strategy. Guardrail metrics will be monitored; significant negative shifts in guardrails will prompt further investigation regardless of primary metric results.