Incrementality testing for ad campaigns
July 18, 2026 · 6 min read
Incrementality testing is an experiment in which you deliberately stop showing ads to a comparable share of your market, run the campaign to everyone else, and measure the difference in outcomes between the two groups. That difference is the only figure in advertising produced by a control group rather than a reporting rule, which is why no dashboard can supply it.
Attribution works the other way round. It decides which touchpoint deserves credit for conversions that already happened, a useful exercise with its own depth in multi touch attribution. But no credit model can tell you whether a purchase would have happened without the ad. Only a group that never saw the ad can.
What a holdout costs before it tells you anything
Start with the honest part. A holdout is deliberately unspent budget and unserved demand. Withhold ads from a fifth of your market for a month and you accept whatever revenue that fifth would have produced above its organic baseline. The cost lands before any insight does, which is why most accounts never run one.
It is worth paying when the decision is large and expensive to get wrong: whether a retargeting campaign reporting an excellent return is actually creating sales, or whether branded search is buying clicks you would have received anyway. If the question is which headline performs better, run a creative test instead. Holdouts returned to fashion largely because platform reported conversions became less complete after the privacy changes described in the iOS 14 shift and Facebook ads.
Three ways to build the control group
User level splits inside the platform
Meta and Google both offer experiment products that randomly assign people to a treated group and a holdout, then report the gap. This is the cleanest design available, because randomisation happens at the person level and the platform observes both sides, and also the most constrained, because eligibility, spend requirements and supported events change over time. Check the current official documentation for Conversion Lift and Google Ads experiments rather than any blog post from last year, this one included.
Geographic splits
You divide the country into regions, keep the campaign running in some and switch it off in others, then compare totals. Geo splits earn their place when the platform will not give you a user level test, and when a meaningful share of orders arrive by phone or in store, since you compare regional totals rather than tracked events.
On and off time tests
Two weeks on, two weeks off, compare. Easiest to run and weakest to trust, because the periods differ in more than the ads: payday timing, weather, a competitor’s push, a public holiday. Repeat the cycle several times or treat the result as an anecdote.
In a geo split, matching the regions is most of the work
The usual failure is choosing regions because they are convenient rather than because they behave alike. Before launching, pull at least eight weeks of history per candidate region and study the ratio between them, not the absolute numbers. Two regions make a usable pair when the ratio of their weekly orders stays roughly steady; if it swings on its own, the test will produce a difference you cannot interpret.
Turkey makes this harder than many markets, because a single metro can dwarf everything else in the account. A split that puts İstanbul on one side is not a split, it is two different businesses. Leave the largest metro out entirely and run across matched second tier cities, or split by district within the metro if your platform allows it.
Then watch for contamination: people travel, and a national influencer or PR push reaches both sides at once.
A worked example, in purchases and in lira
Suppose you split the addressable audience into two equal groups for one month. The exposed group sees the campaign, the holdout group sees nothing from you on that channel. At month end the exposed group has produced 600 purchases and the holdout group 500.
The incremental result is the difference: 600 − 500 = 100 purchases. That is what the ads created. The other 500 arrived on both sides regardless, which means five of every six purchases in the exposed group would have happened without the campaign.
Now put money against it. Say the campaign spent ₺30,000 reaching the exposed group. Divide that across every purchase in that group and you get a comfortable looking 30,000 ÷ 600 = 50 lira per purchase. Divide it across the purchases the campaign actually created and you get 30,000 ÷ 100 = 300 lira. The incremental cost is 300 ÷ 50 = 6 times the reported one, and it is the figure that should drive the budget decision.
A lift test is therefore read at the blended level, not inside campaign reporting: you are comparing two populations’ totals, so total revenue against total spend is the correct frame, and the MER calculator holds that shape. No dashboard can show the holdout, because the holdout has no ads to report on.
How long incrementality testing has to run
The binding constraint is not calendar time, it is conversion volume, and you can estimate it before launching. For counted events, ordinary random variation is roughly the square root of the count. In the example above the two groups produced 600 + 500 = 1,100 purchases between them, and the square root of 1,100 is about 33. An observed difference of 100 is around three times that, so it is unlikely to be an accident of timing.
Now shrink the same test to a shop doing a tenth of that volume: 60 purchases in the exposed group, 50 in the holdout. The gap sits at the same proportion, but in absolute terms it is only 10 purchases. Across both groups you have 60 + 50 = 110 events, and the square root of 110 is about 10 as well. Signal and noise are the same size, and that test cannot answer the question however carefully it was built.
Because variation shrinks with the square root of the count, you need roughly four times the events to halve it relative to the total. At fixed traffic that means four times the calendar: a muddy one month test becomes a four month test, and few businesses will hold a holdout that long. Work it out before you launch. If the volume cannot produce a readable answer, do not spend the holdout at all.
Inconclusive is not the same as negative
This is where most lift tests are misread. A test reporting no distinguishable difference has not shown that the channel does nothing; it has shown that the channel’s effect is smaller than what this test could resolve. Those are different findings, and only one justifies switching a campaign off.
Read the interval, not the point estimate. If the result is compatible with anything from a small negative effect to a large positive one, the instrument had no resolution. Ask instead whether the holdout was too small, the window too short, the spillover too heavy, or a seasonal event landed in the middle.
A genuinely negative result looks different: a clear gap in the wrong direction, repeated, usually on a campaign reaching people already on their way to buying. That is a real finding. An inconclusive one is an instruction to design a bigger test, or to accept that this decision will be made by judgement.
Before you launch the test
- Write down the decision the test will settle and the action you will take for each outcome, before any data arrives.
- Pick one conversion event and one time window, and hold both to the last day. Changing either mid test destroys the comparison.
- Estimate expected volume in the smaller group and check that the difference you care about is larger than the square root of the combined count.
- Freeze other marketing changes in the test regions for the duration.
- Log anything you could not control, such as a stockout or supplier delay, as it happens.
A lift result will not repair your tracking, but it gives you an independent check on it, which is more than another attribution model offers. And when a test says a channel is genuinely additive, the next problem is raising spend without breaking the delivery that produced it, which is where automatic scaling belongs.