All Posts
Measurement

Incrementality Testing on a $50k Budget: The Geo-Holdout Design

Platform lift studies lock out mid-market advertisers on conversion volume, not spend. Geo holdouts don't care how many conversions you have. Here's how to design one you can actually defend in a budget meeting.

· 5 min read
RevenueProven Team
By RevenueProven Team· Editorial
Line and bar performance analytics charts displayed on a laptop screen

Most teams reach for incrementality testing and immediately hit the same wall: the platform-native lift study they wanted to run won't accept them. Meta's self-serve Conversion Lift, for example, expects an ad account with a campaign in the past year carrying at least $5,000 in spend (Meta Business Help Center) plus 500 optimized conversions and a Conversions API setup with an Event Match Quality score above 5. Plenty of mid-market B2B and considered-purchase advertisers clear the spend bar and fail the conversion-volume bar — a long sales cycle simply doesn't produce 500 purchase events in a testable window.

The reflex at that point is to give up on causality and go back to arguing about last-click versus data-driven attribution in the platform UI. That's the wrong move. There is a well-documented alternative that doesn't care about your conversion volume per user, doesn't need a user-level identity graph, and has been running inside Google since before the privacy wave: the geo holdout.

Why geo works when user-level testing doesn't

A geo experiment ignores individual users entirely. You split non-overlapping geographic regions into treatment and control, use geo-targeting to actually withhold ads from control, and measure the difference in your business outcome at the region level. Google researchers Jon Vaver and Jim Koehler formalized this design in a 2011 paper (Measuring Ad Effectiveness Using Geo Experiments), and a later Google paper extended it into a time-based regression framework specifically for advertisers who have too few geographic units for the original geo-based regression to work.

That extension is the part that matters for mid-budget advertisers. The classic design assumes you have many comparable geos to randomize across. If you sell into eight metros, you don't. The time-based approach borrows statistical power from the pre-period instead of from the number of regions — which is exactly the trade you need to make when your footprint is small.

Meta's open-source GeoLift library takes the same problem from another angle, using synthetic control methods to build a weighted composite of untreated regions that mimics your treatment region's pre-period behavior. Notably, GeoLift ships power calculators and a market-selection function that tells you the minimum detectable effect and minimum budget required before you spend anything. Read that in reverse: it is a pre-mortem. If the calculator says your design can only detect an enormous effect, you have just saved yourself six weeks of running a test that was always going to come back "inconclusive."

Designing a holdout you can actually defend

Four decisions determine whether your result is usable.

Pick the outcome at the region level, not the user level. Revenue, qualified pipeline created, or signups by billing region. If you can't attribute the outcome to a geography without a tracking cookie, you can't run this test. For B2B, billing address or company HQ region on closed-won records usually works; form-fill IP geo is an acceptable proxy for top-of-funnel.

Withhold hard, not softly. The most common way these tests die is a leaky holdout — brand search, retargeting, an always-on nurture sequence, or a partner campaign still running in the control region. Kill every channel you're testing in control, or you're measuring a diluted effect and will read it as "no lift."

Change exactly one thing. Geo tests are blunt instruments. If you simultaneously refresh creative, change bidding, and shift budget, the result is uninterpretable no matter how clean the statistics are.

Set the window from the sales cycle, not the calendar. A four-week test on a nine-week sales cycle measures ad-to-lead, not ad-to-revenue. That's a legitimate test — just say so up front and pick a leading indicator you trust, rather than quietly relabeling it as a revenue result later.

Feeding the result somewhere it compounds

A single geo test answers one question about one channel in one period. Its real value is as a calibration anchor. Google's open-source MMM, Meridian, is explicitly built to take incrementality experiment results as priors, so the model can be calibrated against real-world causal results rather than correlational spend curves (Google Ads blog, January 2025).

That's the stack worth building toward: platform attribution for day-to-day optimization decisions, geo experiments for periodic ground truth, and a mix model that inherits those experiments as priors so the ground truth doesn't evaporate the week after the test ends. Multi-touch attribution stays in the picture as an in-flight steering signal — not as the number you take to the board. If your funnel model itself is contested, fix that first; we covered the committee-driven rebuild in The B2B Buying Committee Broke Your Funnel.

What to do this week

  1. Run GeoLift's power calculator against twelve months of region-level outcome data before designing anything. If minimum detectable effect comes back implausibly large, your problem is data granularity, not measurement philosophy.
  2. Audit for holdout leakage: list every channel, retargeting audience, and partner campaign that can reach your control regions, and write down how each one gets suppressed.
  3. Pick one channel — the one with the loudest internal disagreement about its value — and scope a single-variable test with a window matched to your sales cycle.
  4. Decide now what you will do at each outcome. A test with no pre-committed decision rule becomes a debate, and debates default to the status quo.