The short version: Advantage+ and Performance Max can report strong ROAS while counting buyers who would have purchased anyway. A controlled holdout estimates the revenue ads caused beyond that baseline. Test before major budget decisions. Repeat when spend, channel mix, or demand changes.
The best-reported campaign can still mislead you. It shows a 9x return, so you add budget, yet blended ROAS barely moves. The campaign may be claiming sales that were already likely.
AI-automated buying makes this gap easier to miss. Advantage+ and Performance Max optimize attributed outcomes, not a causal holdout result. They may reach buyers who are already close to purchase. This post explains how our Paid Media work separates reported results from estimated lift.
What Is Incrementality Testing?
Incrementality testing estimates the revenue your ads caused beyond what would have happened without them. One group sees ads while a comparable control group does not. The difference in outcomes is incremental lift, subject to the test’s confidence range.
The gap can change a budget decision. Standard ROAS follows the platform’s attribution settings. A holdout asks a different question. Common Thread Collective’s client analysis put median incremental ROAS for Google branded search at 0.27x in its sample. That finding should not be treated as a universal branded-search benchmark.
Google’s Conversion Lift documentation distinguishes attributed conversions from causal, incremental conversions. At the P&L level, the test asks: how much revenue would likely disappear if this campaign stopped?
Why Do AI-Automated Campaigns Inflate Reported ROAS?
AI campaigns can overstate incremental value because platform optimization and business measurement answer different questions. The system seeks attributed conversions or value. You need to know how much demand the campaign added beyond the control baseline.
Advantage+ can widen this gap compared with manual campaigns. In Haus’s incrementality dataset, 58% of brands had higher incremental ROI from manual campaigns. The difference between reported and incremental performance was about 12 percentage points in that sample.
That does not mean Advantage+ should be off. It may still win your test. But reported ROAS alone cannot rank campaigns by causal value. Reported and incremental ROAS are different measurements.
| Metric | Reported ROAS | Incremental ROAS |
|---|---|---|
| What it counts | Every conversion the platform attributes | Only sales the ad actually caused |
| Who reports it | The ad platform | The experiment design and analysis |
| Main limitation | Depends on attribution rules | Depends on power and control quality |
| Effect of AI automation | Can favor attributed demand | Estimates lift against a baseline |
| Use it to | Optimize inside a campaign | Decide budget between campaigns |
What Are the Main Incrementality Test Methods?
Three methods matter for most advertisers: geo holdouts, platform conversion lift, and public service announcement (PSA) tests. Each creates a control in a different way. The right method depends on geography, conversion volume, platform access, and the decision you need to make.
A geo holdout splits markets into test and control groups, then compares revenue. Conversion lift uses a platform-managed, user-level holdout inside Meta or Google. A PSA test shows the control group a charity ad. This separates ad exposure from audience quality.
| Method | Best for | Main limitation |
|---|---|---|
| Geo holdout | Automated campaigns, cross-channel truth | Needs national or multi-region spend |
| Conversion lift | Single-platform reads, faster setup | Graded by the platform being tested |
| PSA test | Isolating exposure from audience quality | Costs budget on the control ad |
For cross-channel or publisher-independent questions, we often lean on geo holdouts. Google’s geo-experiment research describes how randomly assigned regions can measure advertising effects. A platform lift study can be better when user-level randomization and sufficient power are available.
Not sure how much of your ROAS is real?
We build incrementality into how we run paid media, so budget decisions rest on caused revenue, not platform claims.
Book a Free Strategy CallHow Do You Run a Geo Holdout Test on an AI Campaign?
Assign comparable markets to test and control, then hold the campaign dark in the control group. Keep other conditions as stable as possible. Compare outcomes only after the design reaches its planned duration and statistical standard. The estimated difference is the incremental result.
Careful setup makes the number trustworthy. Match your test and control markets on size, seasonality, and baseline demand before you start. A common failure is comparing your best market against your weakest and calling the difference lift.
Write the exclusion rules and decision threshold before launch. Otherwise, a noisy result can become whichever story the team prefers.
Document planned spend, conversion lag, and outside promotions too. A pricing change or major email campaign can break market comparability. Flag those events before analysis and explain whether the test can still support a decision.
Run these steps in order:
- Use enough markets per group for the planned statistical power. Our starting screen is often 10 or more.
- Match the groups on historical revenue and trend, not just population.
- Establish a clean pre-period baseline before any change.
- Hold the control dark for the full planned window. Our tests often run 4 to 6 weeks.
- Compare incremental revenue against incremental spend to get true incremental ROAS.
Then act on the gap. A campaign may report 8x but test at 2.8x. It can remain profitable while supporting much less budget. That correction reshapes the plan. We apply the same discipline during a full paid media audit.
How Often Should You Re-Test Incrementality?
Re-test around major decisions, such as a new campaign type or a large budget shift. Incrementality is not fixed. It moves with auction competition, creative, spend, and demand mix. Google says most advertisers run one or two Conversion Lift studies per year. Higher-spend teams may justify quarterly testing.
Treat the lift number as a financial control, not a one-time audit. A campaign that tested at 3.5x in Q1 can drift as spend grows. Marginal buyers may add less incremental value. Pair that result with the unit economics in our CAC benchmarks work. Together, they show where the next dollar belongs.
The test gives you a better budget map. You stop scaling campaigns that harvest demand and start funding the ones that create it. In an account full of AI automation, that distinction determines which campaigns deserve more budget.
Want to estimate how much paid media is incremental? Book a Free Strategy Call, and we will map the right test.
Like this? Get the next one.
Short emails. New posts as they ship.