There is no statistically defensible universal rule such as "spend $50 per creative" or "run every ad for three days". A test budget is an output, not a policy, and it depends on five inputs:
- 1.The outcome being tested
- 2.Its baseline rate
- 3.The smallest improvement worth acting on
- 4.The acceptable false-positive risk
- 5.The cost of obtaining each observation
Small relative lifts at low baseline rates require very large samples. Most modest Meta budgets cannot prove conversion-rate differences between individual creatives, which is fine — they can still learn. What has to change is the readout, so it matches the evidence the budget actually bought.
The four statistical inputs
Baseline rate is what the metric does today. If the current link click-through rate is 1.0%, that is the baseline; if the landing-page conversion rate is 3.0%, use 3.0%.
Minimum detectable effect (MDE) is the smallest change worth acting on. A lift from 1.0% to 1.2% is a 0.2 percentage-point absolute change and a 20% relative lift. Both framings matter: the absolute gap drives the sample size, the relative one drives whether anyone cares.
Significance level is commonly 5%, the 95% confidence framework. It controls the chance of declaring a difference when none exists, under the model’s assumptions.
Statistical power is commonly 80%: the probability of detecting the pre-declared effect if it is genuinely present. Skipping this input is the most common planning error, because a test can be built that could never have found the effect it was looking for.
NIST’s sample-size guidance shows why baseline proportion, effect size, significance, and power all belong in one calculation. Everything below uses the common normal approximation for two independent proportions at a two-sided 5% significance level and 80% power. See NIST/SEMATECH: sample sizes required for proportions.
The two-proportion approximation
For two equal-sized cells, the approximate sample required per cell is:
n = [ 1.96 × √(2p̄(1−p̄)) + 0.842 × √(p1(1−p1) + p2(1−p2)) ]² ÷ (p2 − p1)²
Where p1 is the baseline rate, p2 is the target rate, p̄ is the average of p1 and p2, 1.96 is the z-value for a two-sided 5% significance level, and 0.842 is the z-value for 80% power. For rare events or small samples, use an exact or simulation-based method with a statistician. This is a planning tool, not a verdict.
Worked sample-size and budget scenarios
Every row below is a worked model using the formula above, not a forecast and not an AdRiseLab customer result.
| Outcome | Baseline to target | Observations per cell | Total observations | Cost assumption | Modeled test budget |
|---|---|---|---|---|---|
| Link CTR | 1.0% to 1.2% | 42,694 impressions | 85,388 impressions | $15 CPM | $1,281 |
| Link CTR | 2.0% to 2.4% | 21,109 impressions | 42,218 impressions | $15 CPM | $633 |
| Site CVR | 3.0% to 3.6% | 13,914 clicks | 27,828 clicks | $2 CPC | $55,656 |
| Site CVR | 5.0% to 6.0% | 8,158 clicks | 16,316 clicks | $2 CPC | $32,632 |
| Site CVR | 3.0% to 4.5% | 2,518 clicks | 5,036 clicks | $2 CPC | $10,072 |
Two comparisons in that table are worth sitting with. The same 20% relative lift costs about $1,281 to detect on link CTR and about $55,656 to detect on site conversion rate, because clicks cost roughly 130 times what impressions cost. And widening the MDE on the same metric — 3.0% to 4.5% instead of 3.0% to 3.6% — cuts the required sample by 82% and the budget from $55,656 to $10,072.
That second row is the practical lever. A small advertiser can usually evaluate attention long before it can afford to prove downstream revenue differences, and testing a genuinely bigger creative difference is far cheaper than testing a small one more patiently. Actual CPM, CPC, event rates, dependence between observations, and delivery imbalance will all move these numbers.
Why equal budget per ad may still not be an A/B test
An auction platform may not distribute impressions equally across ads placed in the same ad set. When delivery is optimised dynamically, one creative can absorb most of the spend. That is useful for campaign performance and it does not create a controlled experiment.
A defensible comparison needs all six of these, and the sixth is the one teams skip:
- A pre-declared primary outcome
- Comparable audiences and dates
- Controlled differences between cells
- Enough observations in each cell
- No unlogged landing-page, offer, or tracking change
- A stopping rule chosen before anyone looks at the result
Meta Blueprint teaches the formal A/B and lift-test concepts, and the platform’s own experiment tooling is the right instrument when causal evidence matters more than routine optimisation. See Meta Blueprint: Ad Creative Testing.
Match the metric to the decision
| Decision | Better primary outcome | Why |
|---|---|---|
| Does the opening stop attention? | Three-second view or hook rate, with placement context | Fast signal, but not a revenue metric |
| Does the message earn a click? | Link CTR or landing-page views per impression | Tests message-to-destination interest |
| Does the click convert? | Landing-page CVR | Separates traffic quality from page performance |
| Does the ad create profitable customers? | Contribution-adjusted CAC or incremental profit | Closest to business value, slowest to establish |
Do not promote a high-CTR ad to winner status if its clicks do not convert, and do not kill a lower-CTR ad that produces more qualified or higher-value customers. The metric you can afford to measure and the metric that decides the business are frequently not the same one, and the gap between them should be stated in the readout rather than quietly ignored.
A budget worksheet to fill in before launch
| Input | Your value |
|---|---|
| Primary outcome | |
| Baseline rate | |
| Minimum useful target rate | |
| Relative lift | |
| Significance level | 5% unless deliberately changed |
| Power | 80% unless deliberately changed |
| Required observations per cell | |
| Expected CPM or CPC | |
| Number of cells | |
| Modeled total budget | |
| Maximum test duration | |
| Decision after a positive result | |
| Decision after an inconclusive result |
The last two rows exist because an unplanned inconclusive result is where most testing programmes quietly break. Decide in advance what happens when nothing separates, and the test stops being a coin flip you interpret afterwards. The creative volume planner sizes the concept pipeline this worksheet consumes, and the Meta ads ROI calculator converts a modeled lift into the revenue it would have to produce to be worth funding.
When the required budget is too high
Six options, roughly in order of how often they are the right one:
- 1.Test a larger creative difference. A different proposition or proof mechanism is far more likely to produce a detectable effect than a minor visual edit, and it costs the same to run.
- 2.Use a faster upstream metric, carefully. Treat it as directional evidence about attention, not proof of profit.
- 3.Reduce the number of cells. Five underpowered variants are not better than one clean comparison.
- 4.Pool only when the unit is genuinely comparable. Merging unrelated countries, placements, or funnel stages to inflate the sample buys a number, not evidence.
- 5.Run longer without repeatedly peeking and stopping. Opportunistic checking inflates false-positive risk well beyond the nominal 5%.
- 6.Apply business thresholds first. If an effect is too small to matter economically, statistical significance does not rescue it.
Option one deserves the emphasis it gets. The arithmetic in the scenarios table says a bigger MDE is the cheapest variable in the whole calculation, and creative difference is the only input on that list you fully control.
Related Reading
For the operational layer around these numbers, see the Meta creative testing guide and how many ad creatives a Meta account needs. The Meta ads A/B testing guide covers the setup mechanics inside Ads Manager, and a 12-point Meta ads audit is the faster move when the account has measurement problems that would corrupt any test you run.
