Reader question: If version B gets a lower cost per conversion than version A, what exactly have we learned—and what remains unproven about the campaign?
The answer is in the comparison. An A/B creative test puts a current ad and a changed ad in comparable delivery conditions, then asks which performs better. Both groups see ads, so the result can guide the next edit but does not show that either created additional demand.
Relative performance is the gap between two exposed versions. Incrementality is the outcome attributable to exposure compared with a credible no-exposure counterfactual. Put that distinction in the brief first.
Name the question before naming the winner
Write the estimand—the quantity the test is meant to describe—in ordinary language. An A/B question asks which ad produces the better rate among assigned people. An incrementality question asks what additional outcomes occurred because similar people saw the campaign rather than not.
Google Ads separates these ideas: Experiments compare settings, ad copy, bidding, or assets; Lift Studies compare exposed treatment with unexposed control for incremental impact. Google Ads Experiment Center
| Question | Comparison | Safe first sentence in the readout |
|---|---|---|
| Which creative performs better? | Current versus changed ad, fixed skeleton | “B differed from A under this setup.” |
| What did the campaign add? | Exposed group versus credible no-ad holdout | “The study estimates a difference, subject to design and measurement limits.” |
Diagram in words: two exposed arms, one deliberate change
For a proposed creative test, picture the flow like this:
Eligible traffic
↓ assignment supplied by the platform or test design
Control: current creative + fixed campaign setup
Treatment: one changed creative element + the same fixed setup
↓
Compare the predeclared metric; describe the result only within this comparison.
The example is illustrative; no campaign or result is reported. A team might compare a creator edit opening with a product demonstration against one opening with a use-case sentence. Keep eligibility, objective, placement, bidding, budget, destination, offer, dates, and conversion definition fixed. If several change, the result no longer isolates the opening.
Google recommends one variable and one or two metrics chosen before launch, with no base-campaign changes that muddy comparison. Google Ads, “Test with confidence”
Hold the skeleton still, including the measurement
“Same audience” is not enough. Preserve exclusions, objective, placements, optimization event, bid rules, budget, schedule, destination, offer, approvals, measurement window, and event definition. A click, view, lead, or purchase is not interchangeable.
Record assignment. Google says its split is applied to eligible auctions before targeting; cookie- and search-based methods can expose people to one arm or both across searches. Google Ads, “The statistical methodology behind experiments” Preserve the exact method and ask whether arms had comparable opportunity.
If supported, an A/A check can reveal imbalance first. Google defines A/A arms as identical throughout, including settings and approvals; changes should reach both at once. Google Ads, “The statistical methodology behind experiments”
Read a winner as a conditional comparison
If B produces a better predeclared metric, say only that B performed better than A under recorded conditions. That does not establish preference, portability, or incremental conversions. A statistically distinguishable result can still be too small or expensive to roll out.
Google says its reporting uses bucketed data, jackknife resampling, and two-tailed testing with a 95% confidence interval. Google Ads, “The statistical methodology behind experiments” That is not universal. Report estimate, uncertainty, metric, and decision separately; “winner” is not ROI.
Incrementality requires a no-ad counterfactual
Incrementality needs a group representing what would have happened without ad exposure. Google describes Lift Studies as treatment shown ads versus control not shown ads, with user- and geo-based approaches depending on campaign and goal. Google Ads, “About lift studies” Verify availability and measurement gaps.
A proposed three-arm design makes the two questions visible:
Holdout: no campaign exposure ──┐
Creative A: current ad ─────────┼─ compare A and B for creative performance
Creative B: changed ad ────────┘ compare exposed arms with holdout for campaign effect
This is a planning diagram, not an executed study. B versus A asks which exposed treatment is better; exposed arms versus holdout ask a causal question that may require a lift product or geographic design. If a platform reports only A versus B, label it a creative comparison and leave the no-ad claim open.
Build the readout in layers
An honest decision record can fit on one page:
- Hypothesis: the single creative change and why it might affect the chosen outcome.
- Assignment: eligible population, split method, arms, dates, and delivery caveat.
- Fixed inputs: campaign, audience, destination, offer, and measurement settings.
- Outcome: each arm’s delivery, spend, rate, count, and platform uncertainty; use predeclared metrics.
- Interpretation: relative result, practical tradeoff, and what it cannot answer about no-ad impact.
- Decision: ship, iterate, hold, or design a separate incrementality study; name missing evidence.
Stop when creative changed with targeting or bidding, arms lacked comparable opportunity, the metric was chosen after the fact, or no credible holdout exists for a causal claim. A/B testing remains useful; keep its promise small enough to be true.
Sources and limitations
- Google Ads, “About the Google Ads Experiment Center” — checked September 4, 2026; supports the documented distinction between Experiments and Lift Studies, including settings/ad-copy comparisons and exposed-versus-unexposed group methodology. Limitation: this is Google Ads product guidance; availability, eligibility, and implementation details vary by account and platform, and the documentation is not an independent performance study.
- Google Ads, “Test with confidence with the Experiments page” — checked September 4, 2026; supports setting a hypothesis, testing one variable at a time, preselecting one or two success metrics, avoiding base-campaign changes, and keeping experiment records. Limitation: it is operational guidance for Google Ads, not a guarantee that any test will be powered or causal.
- Google Ads, “The statistical methodology behind experiments” — checked September 4, 2026; supports Google’s documented split behavior, cookie/search assignment distinction, A/A conditions, and its jackknife/95% confidence-interval reporting method. Limitation: these are Google’s platform-specific methods; they should not be generalized to another platform or treated as a universal significance rule.
- Google Ads, “About lift studies” — checked September 4, 2026; supports the treatment/no-ad control distinction and the documented user-based and geo-based conversion-lift approaches. Limitation: eligibility, budget, measurement gaps, and study design determine whether a lift method is available or fit for a specific campaign.