Incrementality Test Calculator
Before you hold ads back from part of your market, find out what the test can actually see. Enter your conversion volume, duration, and holdout share to get the minimum detectable lift, how long you'd need to run to catch the effect you expect, and a straight answer to the question that matters most: if this comes back flat, does that mean anything?
Minimum detectable lift
15.3%
The smallest true effect this test can distinguish from noise
Conversions in window
1,600
480 holdout · 1,120 exposed
Weeks to detect your lift
9.3
To reach 10.0% MDE at this volume
MDE at an even 50/50 split
14.0%
A 50% holdout is always the most sensitive — and the most expensive
Underpowered — a null result would be inconclusive, not negative.
You expect 10.0% but this test can only resolve effects of 15.3%or larger. If it comes back flat you will not be able to tell “the ads don't work” from “the test couldn't see it” — the single most expensive mistake in incrementality testing. Run it for 9.3 weeks, raise the holdout toward 50%, or pick a higher-volume conversion to measure.
- Holdout conversions (control)
- 480
- Exposed conversions (test)
- 1,120
- Cost of your split vs. 50/50
- +1.3% MDE
Conversions are modelled as Poisson counts, the usual approximation for holdout tests measured on conversion volume. It assumes the groups are genuinely comparable before the test — matched-market selection and a pre-period parallel-trends check do more for validity than any sample-size number. Everything is computed in your browser.
The mistake this exists to prevent
Incrementality testing is the honest answer to “is this spend actually doing anything?” — you withhold ads from a comparable slice of the market and compare. The logic is sound. The failure is almost always arithmetic rather than conceptual.
A test runs for two weeks on a few hundred conversions, comes back flat, and gets reported as evidence that the channel doesn't work. But at that volume the smallest resolvable effect might be 25–30%. A genuine 12% incremental lift — a very good result — is invisible. The test didn't measure the ads; it measured its own noise floor, and a profitable channel gets cut on the strength of it.
The fix costs nothing: calculate the minimum detectable lift beforelaunching. If the MDE is larger than any lift you could plausibly get, the test cannot succeed and shouldn't run in that shape. Change the duration, widen the holdout, or measure a higher-volume conversion — but decide that up front, not after you've spent three weeks buying an unreadable answer.
Two things this deliberately doesn't claim. Power is not validity: comparable groups and a parallel pre-period matter more than any sample-size number, and a well-powered test on mismatched geos will give you a confident wrong answer. And significance is not profitability — a statistically real 3% lift can still lose money.
Frequently asked questions
What is minimum detectable lift (MDE)?
The smallest true effect your test can reliably distinguish from random noise, given your conversion volume, holdout split, confidence level, and power. If your MDE is 30% and the real incremental lift is 10%, your test will almost certainly come back flat — not because the ads did nothing, but because the experiment was never sensitive enough to see it.
Why does an inconclusive test get read as a negative result?
Because a flat readout looks like proof that spend isn't working. It isn't. Without an MDE calculated up front you cannot separate 'no effect' from 'no resolution', and the usual consequence is cutting a channel that was actually profitable. Computing MDE before you launch is what makes a null result interpretable.
How big should my holdout be?
A 50/50 split is mathematically the most sensitive — MDE is minimised when both groups are the same size. It is also the most expensive, because you are withholding ads from half your market. The calculator shows what your chosen split costs you in MDE versus an even one, so you can trade sensitivity against lost revenue deliberately.
How many conversions do I need for an incrementality test?
It scales with the square of the effect you're chasing: halving the lift you want to detect needs roughly four times the conversions. As a rough anchor at 95% confidence and 80% power with an even split, around 1,000 conversions a week gets you to roughly a 10% MDE within a few weeks, while 100 a week struggles to resolve anything under about 25%.
Does a good MDE mean my test is valid?
No — it means your test is adequately powered, which is a different thing. Validity comes from the groups being genuinely comparable: matched-market selection, a pre-period check that the two groups' trends ran parallel, and no other changes launching mid-test. A perfectly powered test on badly matched geos produces a confident wrong answer.
Does this calculator send my numbers anywhere?
No. Everything is computed in your browser — nothing you type is transmitted or stored.
Designing the test itself
Sizing is the easy half. Choosing matched markets, validating the pre-period, and making sure the conversions you're measuring are tracked correctly in the first place is where holdout tests usually go wrong.