Skip to content
Paid Media

Incrementality Testing for AI Ad Campaigns

By Alex Montas Hernandez
Incrementality Testing for AI Ad Campaigns

Incrementality testing estimates how many sales your ads add beyond what would have happened without them. It compares a group exposed to ads with a similar group that does not see them. That helps answer a different question from a platform report, which assigns credit under its attribution rules.

The short version: Advantage+ and Performance Max may count sales from people already likely to buy. A controlled holdout helps estimate the additional sales caused by advertising, within the limits of the test. Use it before major budget changes and repeat when spending, channels, or demand changes.

This guide compares the test methods, explains what a reliable comparison needs, and shows how results can change a budget decision. It is part of how our Paid Media work checks reported performance against business outcomes.

What Is Incrementality Testing?

Incrementality testing estimates the revenue your ads caused beyond what would have happened without them. One group sees ads while a comparable control group does not. The difference in outcomes is incremental lift, subject to the test’s confidence range.

That distinction can change a budget decision. Standard ROAS follows the platform’s attribution settings, while a holdout estimates what happened because the ads ran. Common Thread Collective’s client analysis put median incremental ROAS for Google branded search at 0.27x in its sample. That finding should not be treated as a universal branded-search benchmark.

Google’s Conversion Lift documentation distinguishes attributed conversions from causal, incremental conversions. At the P&L level, the test asks: how much revenue would likely disappear if this campaign stopped?

Why Do AI-Automated Campaigns Inflate Reported ROAS?

AI campaigns can overstate incremental value because platform optimization and business measurement answer different questions. The system seeks attributed conversions or value. You need to know how much demand the campaign added beyond the control baseline.

Advantage+ can widen this gap compared with manual campaigns. In Haus’s incrementality dataset, 58% of brands had higher incremental ROI from manual campaigns. The difference between reported and incremental performance was about 12 percentage points in that sample.

Advantage+ may still win your own test. The finding is a reason to check, rather than switch it off: reported ROAS and incremental ROAS measure different things, so one cannot stand in for the other.

MetricReported ROASIncremental ROAS
What it countsEvery conversion the platform attributesOnly sales the ad actually caused
Who reports itThe ad platformThe experiment design and analysis
Main limitationDepends on attribution rulesDepends on power and control quality
Effect of AI automationCan favor attributed demandEstimates lift against a baseline
Use it toOptimize inside a campaignDecide budget between campaigns

What Are the Main Incrementality Test Methods?

Three methods matter for most advertisers: geo holdouts, platform conversion lift, and public service announcement (PSA) tests. Each creates a control in a different way. The right method depends on geography, conversion volume, platform access, and the decision you need to make.

A geo holdout compares revenue across test and control markets. Conversion lift instead uses a platform-managed holdout of users inside Meta or Google. A PSA test shows the control group a charity ad to separate the effect of ad exposure from audience quality.

MethodBest forMain limitation
Geo holdoutAutomated campaigns, cross-channel truthNeeds national or multi-region spend
Conversion liftSingle-platform reads, faster setupGraded by the platform being tested
PSA testIsolating exposure from audience qualityCosts budget on the control ad

For cross-channel or publisher-independent questions, we often lean on geo holdouts. Google’s geo-experiment research describes how randomly assigned regions can measure advertising effects. A platform lift study can be better when user-level randomization and sufficient power are available.

Not sure how much of your ROAS is real?

We build incrementality into how we run paid media, so budget decisions rest on caused revenue, not platform claims.

Book a Free Strategy Call

How Do You Run a Geo Holdout Test on an AI Campaign?

Assign comparable markets to test and control groups, with no campaign ads shown in the control markets. Keep other conditions stable. Once the test meets its planned duration and statistical standard, compare outcomes to estimate the incremental result.

Careful setup makes the number trustworthy. Match your test and control markets on size, seasonality, and baseline demand before you start. A common failure is comparing your best market against your weakest and calling the difference lift.

Write the exclusion rules and decision threshold before launch. Otherwise, a noisy result can become whichever story the team prefers.

Document planned spend, conversion lag, and outside promotions too. A pricing change or major email campaign can break market comparability. Flag those events before analysis and explain whether the test can still support a decision.

Run these steps in order:

  1. Use enough markets per group for the planned statistical power. Our starting screen is often 10 or more.
  2. Match the groups on historical revenue and trend, not just population.
  3. Establish a clean pre-period baseline before any change.
  4. Hold the control dark for the full planned window. Our tests often run 4 to 6 weeks.
  5. Compare incremental revenue against incremental spend to get true incremental ROAS.

Use the gap to adjust the budget. A campaign that reports 8x but tests at 2.8x may still be profitable, though it can support much less spending than the dashboard suggests. We apply the same discipline during a full paid media audit.

How Often Should You Re-Test Incrementality?

Re-test around major decisions, such as a new campaign type or a large budget shift. Incrementality is not fixed. It moves with auction competition, creative, spend, and demand mix. Google says most advertisers run one or two Conversion Lift studies per year. Higher-spend teams may justify quarterly testing.

Keep using lift to guide financial decisions after the first test. A campaign that measured 3.5x in Q1 can become less incremental as spending grows and the next buyers add less value. Pair that result with the unit economics in our CAC benchmarks work. Together, they show where the next dollar belongs.

The test gives you a better budget map. You stop scaling campaigns that harvest demand and start funding the ones that create it. In an account full of AI automation, that distinction determines which campaigns deserve more budget.

Before your next large budget change, we can help you decide whether a holdout would answer the question that matters. Bring your campaign reports, conversion volume, and planned spend change. Book a Free Strategy Call to discuss a test design and the evidence you would need to act on it.

Like this? Get the next one.

Short emails. New posts as they ship.

A
Alex Montas Hernandez

Founder

Previously led growth at TubeBuddy (acquired by BENlabs), scaled Bloomberg's first DTC subscription, and drove measurable growth for brands like Verizon, Samsung, and Intel.

Frequently Asked Questions

What is incrementality testing in paid media?

Incrementality testing estimates the revenue your ads caused, beyond standard attributed reporting. A controlled test withholds ads from one group and exposes a comparable group. The measured difference is incremental lift. User-level conversion lift and well-designed geo experiments are two common methods. Both need enough data and a valid control to support a causal conclusion.

Do Advantage+ and Performance Max over-report ROAS?

They can. Both systems optimize attributed conversions, which may include buyers already close to purchasing. In Haus's dataset, 58% of brands had higher incremental ROI from manual campaigns than Advantage+. The gap was roughly 12 percentage points in that sample. That result is useful evidence, not a universal forecast for every account.

How often should you run an incrementality test?

Test before major budget decisions and repeat when channel mix or demand changes. Google says most advertisers run one or two Conversion Lift studies per year. A high-spend team may test quarterly when it has enough data and can afford the holdout. Lower-volume teams should time tests around the decisions with the greatest financial impact.

Get the next post in your inbox

I write about growth, AI performance creative, and what's actually working in 2026. New posts when I have something real to say.

Or book a strategy call →