Geo Holdout Testing: A Step-by-Step Guide for E-Commerce Brands
By L&Y Decision
Your attribution dashboard tells you which platform claimed each sale. It cannot tell you whether the sale would have happened anyway. A geo holdout test can, and you can run one with a spreadsheet and six weeks of patience.
Why Geographic Markets Are the Right Unit of Measurement
You cannot show ads to a customer and simultaneously not show them ads. That counterfactual problem is why attribution models guess instead of measure. Geography gets around it: pause spend in Denver and Kansas City, keep it running in comparable markets, and the held-out cities become a live control group for your entire marketing mix.
Cookies deprecate, iOS privacy rules tighten, and platforms restrict user-level data a little more each year. Regional revenue data has none of those problems. Your order file has a shipping address on every row, and no platform policy change can take that away from you.
Step 1: Pick the Question Before You Pick the Markets
A geo test answers one question at a time. Write it down in one sentence with a number in it. Good examples: 'Does branded search spend drive any incremental revenue, or does it harvest sales we would get anyway?' or 'What is the true ROAS of Meta prospecting, measured against a holdout?'
The most valuable first test for most DTC brands is branded search, because it is the channel most likely to be claiming credit for demand it did not create. Brands that pause branded search in holdout markets routinely find that 70 to 90% of that 'revenue' shows up anyway through organic listings.
Step 2: Select and Match Your Markets
Split your markets (US states, DMAs, or metro areas) into a treatment group and a control group that behaved alike before the test. Pull 12 months of weekly revenue by region, then pair markets whose revenue curves move together. Denver and Salt Lake City that rise and fall in sync make a good pair; Miami and Minneapolis with opposite seasonality do not.
Practical sizing: Hold out 20 to 30% of revenue-weighted markets. Less than 15% and the signal drowns in noise. More than 40% and the media cost of the experiment starts to hurt. Exclude your one or two mega-markets (usually NY and LA for US brands) from the holdout, since nothing else predicts them well.
Matching check: For each treatment-control pairing, correlation of weekly revenue over the past year should exceed 0.8. If you cannot find pairs that clear that bar, group several small markets together until the aggregates track each other.
Step 3: Decide Direction, Turn Off or Turn On
A holdout (turn spend off in test markets) measures the incrementality of what you already run, and costs you only the sales the channel was driving. A scale-up (turn new spend on) measures opportunity, and costs real budget. Test existing channels with holdouts first; you will usually find waste, and the waste funds the scale-up tests.
Step 4: Set the Runtime and Freeze Everything Else
Run 4 to 8 weeks. You need at least one full purchase-consideration cycle, plus a buffer, because ad effects decay over days or weeks rather than stopping the moment spend stops. Put the end date in writing before launch. Ending a test early because the numbers look good is the most common way brands turn a real experiment back into a guess.
During the test window: no promotions that differ by region, no site redesigns, no new channel launches, no creative overhauls. Anything that changes must change in every market at once, or it contaminates the comparison.
Step 5: Read the Results
For each holdout market, use its matched control markets to project what revenue would have been, then compare against what happened. The simplest defensible method: compute each pair's revenue ratio over the 12 pre-test months, apply that ratio to the control market's in-test revenue, and sum the projections across pairs.
Incremental ROAS: Divide the revenue gap (projected minus actual in holdout markets) by the spend you paused there. That number is your incremental ROAS for the channel. Compare it against the platform-reported ROAS and expect a gap; holdout-measured returns commonly come in at 30 to 60% of what the platform claims.
Sanity checks: Plot the holdout markets against their projections week by week. A real effect appears within the first two weeks and persists. A gap that appears in week five only, or flips sign weekly, is noise. Also check a placebo: run the same math on two control markets, where the 'effect' should be near zero.
What to Do With the Answer
If a channel's incremental ROAS clears your contribution-margin breakeven, scale it and retest at the new spend level in six months, since incrementality falls as spend rises. If it does not clear breakeven, reallocate the budget and run the next test. One test per quarter is a sustainable cadence, and within a year you will have causally measured the channels carrying most of your budget.
The brands that win on paid acquisition in the next five years will not be the ones with the prettiest dashboards. They will be the ones that know, channel by channel, which dollars cause revenue and which dollars take credit for it.
Source
Methodology references: Vaver & Koehler, 'Measuring Ad Effectiveness Using Geo Experiments' (Google Research, 2011); Meta Open Source, GeoLift documentation. L&Y Decision runs matched-market geo tests for e-commerce brands as part of its incrementality measurement practice.
Frequently Asked Questions
What is a geo holdout test?
A geo holdout test measures the causal impact of advertising by turning ads off (or on) in a set of geographic markets while keeping them unchanged everywhere else. You then compare revenue in the held-out markets against what the other markets predict it should have been. The gap is the true incremental effect of the ads.
How long should a geo holdout test run?
Most e-commerce geo tests need 4 to 8 weeks: at least one full purchase cycle plus enough days for the revenue difference to exceed normal market noise. Shorter tests only work for products with same-week purchase decisions and large budgets relative to the markets being tested.
How much revenue do I need to run a geo test?
As a working floor, brands doing roughly $300k+ per month in a single country can usually detect a meaningful effect. Below that, week-to-week noise tends to swamp the signal unless the channel being tested is a large share of total spend.
Do geo tests work for Amazon or marketplace sellers?
They are harder because marketplaces rarely report revenue by buyer region. If you can export regional sales data (Amazon does provide state-level reports), the same design applies. Otherwise, test the channels that drive traffic to your own store, where you control regional measurement.
What is the difference between a geo holdout test and a conversion lift study?
Platform lift studies (Meta, Google) randomize at the user level inside one platform, using the platform's own conversion tracking. Geo holdouts randomize at the market level and read results from your revenue data, so they capture cross-channel effects and don't depend on the platform grading its own homework.
More articles
View all →The Attribution Models Are Getting Smarter. Most Brands Are Still Too Broke to Use Them.
New causal marketing mix models, a 4,200-test CRO study, and fresh CAC benchmarks all say the same thing: sophistication isn't the bottleneck for most brands. Broken tracking, underpowered tests, and blended CAC reporting are.
The Platforms Grading Their Own Homework: Why Your Attribution Data Is Structurally Broken
A peer-reviewed paper from NeurIPS 2025 formally proves what performance marketers have suspected for years: the mechanism that decides which of your ad platforms gets credit for your conversions is mathematically designed to be gamed.
Incrementality Testing 101: What Every E-Commerce CMO Needs to Know
Incrementality is the question every marketing team should be asking: would these customers have converted without our ads? Here's how to find out, without a data science team.
Ready to prove your marketing ROI?
Book a free 30-minute consultation. Bring your ad account numbers; leave knowing which channels earn their budget.
Free. No commitment.