Geo Lift Testing, Explained: A Non-Technical Guide for Marketers

Editorial Policy →

Every ad platform reports its own version of the truth. Meta, Google, and TikTok can each claim the same order, and the gaps keep widening as privacy concerns, privacy laws, cookie deprecation, and app-tracking prompts strip out the user-level data attribution was built on.

In fact, only about 35% of iOS users opt in to tracking when they see Apple's prompt, according to Adjust.

So the measurement problem has shifted. Marketers need to know what their spend actually caused, yet 44% cite accuracy and reliability concerns as their biggest barrier to measuring incrementality.

This guide explains geo lift testing in plain language: what it is, how to design one, what it costs, and how to read the result without a statistics degree.

P.S. If you want help deciding whether geo testing fits your data, our marketing measurement team can assess your regional sales data, design the test, and interpret the results alongside your media strategy.

TL;DR

  • A geo lift test turns an ad channel off, on, or up in some regions and compares sales against similar regions where nothing changed.
  • It answers "did this spend cause the sales?" using regional sales totals, so it needs no cookies, pixels, or user-level tracking.
  • A typical test runs 4 to 6 weeks across 10 to 20+ markets, with at least several months of sales history behind it.
  • It only works when the lift you expect is bigger than the normal week-to-week swing in your sales.
  • The output is incremental return on ad spend (iROAS) with a range around it. The range decides how confident your next budget move can be.

We believe media spend should produce evidence, and spend that teaches you nothing is the real waste. A geo lift test is the cleanest version of that idea. It treats a channel budget as a hypothesis, runs it in the real world, and reports what changed in revenue.

Our team pairs it with creative and channel testing because each answers a different question. The scoreboard we trust is customer acquisition cost and blended revenue, never a platform's own ROAS. The test gives a team the evidence it needs to decide when scaling makes sense.

What Is Geo Lift Testing?

Geo lift testing is an experiment that measures how many sales a marketing channel actually caused. It compares regions where the ads ran against similar regions where they didn't, and the difference between the two is the lift.

The method dates back at least to Google's 2011 paper on measuring ad effectiveness using geo experiments, which proposed the 210 Nielsen media markets in the United States as one ready-made set of geographic units. Because the test reads total sales per region, it doesn't need to know who saw which ad. That's why it survives privacy changes that break user-level tracking.

You'll see the same method under several names:

  • Geo incrementality test, or geographic-based incrementality tests
  • Media geo test, or geo experiment
  • Matched market testing, when one or two test markets are paired with lookalike control markets
  • Geo holdout, when the test means pausing spend in some regions

GeoLift is a separate thing: it's the name of Meta's open-source tool for running these tests. Geo lift testing is one type of incrementality testing, the broader practice of measuring what marketing causes.

Incrementality is the concept (sales that wouldn't have happened without the ads); lift is the measured size of it, usually as a percentage.

How Does a Geo Incrementality Test Work?

A geo incrementality test works like a retail chain trialling a new promotion in a handful of stores. If sales jump in those stores and stay flat in comparable ones, the promotion probably did it. Geo testing applies the same logic to ad spend across geographically similar areas, with statistics doing the comparison.

Every test moves through three periods:

  1. Pre-period. The model studies historical data to learn how your markets normally move together, week by week.
  2. Test period. You change spend in the test markets (the treatment group) and leave everything else alone in the control markets (the control group).
  3. Cooldown. You keep watching after the spend change ends, because some customers buy days or weeks after seeing an ad.

The clever part is the comparison. Few cities behave exactly like each other, so most modern geo-experimental techniques use a synthetic control. The model skips the hunt for one twin city and builds a stand-in from a weighted blend of control markets that tracked the test market closely before launch.

Say Phoenix's synthetic control is 60% Dallas, 30% Denver, and 10% Tampa. If that combination closely matched Phoenix before the test, it can provide an estimate of what Phoenix might have generated without the change in advertising.

Line chart: weekly sales in test markets track their synthetic control during the pre-period, then rise above it during the test period, with the gap labelled as lift.
Sales in the test markets track the synthetic control during the pre-period, then pull ahead once spend changes. The gap is the lift. Source: Fieldtrip original

This is causal inference in practical form: the gap between real sales and the stand-in, after launch, is the lift. The output is a lift figure and an iROAS, each with a range, and the results section below explains how to read them.

Geo Lift Testing vs. Conversion Lift, MMM, A/B Tests, and Attribution

Geo lift testing answers one question better than the other marketing measurement methods: did this channel cause sales, including sales you can't track, like Amazon and retail?

Each of the other methods answers a different question.

MethodQuestion it answersNeeds user-level tracking?Best forMain weakness
Platform attribution (last-click attribution, multi-touch attribution)Which touchpoints came before a conversion?YesDay-to-day optimization inside one platformCounts correlation, and every platform grades its own homework
Platform conversion liftDid this campaign cause conversions among exposed users?Yes, inside the platformSingle-platform campaign checksCan't see off-platform or offline sales
A/B testWhich version of an ad, page, or offer performs better?UsuallyCreative and landing page decisionsDoesn't tell you if the channel itself is incremental
Marketing mix model (MMM)How did each channel contribute over the past year or two?NoLong-term budget mix across all channelsBuilt on correlations in historical data, needs validation
Geo lift testDid changing spend in this channel cause sales, everywhere they happen?NoScale, cut, or launch decisions for one channelNeeds volume, discipline, and weeks of time

Attribution struggles with cross-device leakage, identity resolution, and the general fragility of cookie-based tests and site-based tracking.

One honest caveat. Meta's own GeoLift docs say the company prefers people-based conversion lift where it's feasible, because it has more statistical power. Geo testing wins when you need a read across channels, offline sales, or platforms that don't offer their own lift studies.

Most teams end up running several methods at once. According to the IAB's 2026 State of Data report, 76% of US buy-side decision-makers use incrementality tests and 73% use attribution, yet only 39% use incrementality, attribution, and MMM together.

Model comparison is less about crowning a winner among econometric methods and more about giving each the job it does best, which is where marketing mix modeling fits in.

When Should You Run a Media Geo Test?

Run a media geo test when the lift you expect is bigger than the normal week-to-week swing in your sales. If it isn't, the test can't tell a real effect from noise, however carefully you design it.

For example, suppose a test region normally sells 10,000 units per month and sales commonly fluctuate by around 500 units. If the planned media investment is expected to generate roughly 500 additional purchases, the effect could disappear inside that normal variation.

That is why spend alone does not determine geo-test readiness. What matters is the size of the expected effect relative to the underlying noise in the outcome you are measuring. Market count, sales volume, geographic data quality, purchase cycle, and test duration all affect how much lift a test can reliably detect.

Decision tree with four yes or no questions that decide whether a brand should run a geo lift test.
A four-question decision tree: Can you see sales by region weekly? Is the purchase cycle under a month? Is the expected lift bigger than your weekly sales swing? Can you hold everything else steady for six weeks? Source: Fieldtrip original

Signs a Geo Lift Test Fits

  • You're asking a channel-level question: scale or cut Meta prospecting, check if branded search is incremental, see if connected TV drives store sales.
  • You can see sales by region at least weekly, from your store, Amazon, or retail partners.
  • Customer journeys are short. Measured suggests consideration cycles under a month are good candidates.
  • The channel spends enough to move regional sales activity in a visible way.
  • The channel is hard to measure any other way, like out-of-home advertising, local TV advertising buys, or podcasts, where offline conversions are the whole point.

Signs a Geo Lift Test Doesn't Fit Yet

  • You want to compare two headlines or two creatives. Small tweaks get buried in noise.
  • The purchase takes months, like a car or a mortgage, so the test can't run long enough to catch it.
  • You can't get sales below the national level. Some CPG brands are in this spot, and MMM is the better tool there.
  • Your marketing efforts in the channel are small relative to total sales.
SituationRun a geo lift test?Better option
Deciding whether to cut 30% of Meta prospectingYesn/a
Choosing between two video hooksNoCreative A/B test
Checking if branded search steals organic salesYes, as a holdoutn/a
Sales only visible nationally, monthlyNoMMM
B2B deals with 6-month sales cyclesRarelyPipeline-based holdouts, MMM
Launching a new channel like CTVYes, as a holdoutn/a

How to Design a Geo Lift Test, Step by Step

Designing a geo lift test comes down to three phases: plan the test before launch, run it without interference, then measure it honestly. Most of the work, and most of the ways a test fails, sit in the first phase.

Phase 1: Plan the Test (2 to 4 weeks before launch)

  1. Write the one decision the test will settle. "Should we cut Meta prospecting by 30%?" is a test. "Is Meta working?" isn't. If the question is about which ad performs better, send it to creative testing instead.
  2. Choose the test type. Google's open-source Meridian GeoX names three test designs, and each sets different test spend levels (see the table after these steps).
  3. Pick the KPI and where the data comes from. Use a business outcome the team already trusts, such as orders, revenue, or new customers. Include every place those outcomes can occur, including your website, marketplaces, stores, and retail partners. A test can understate impact if the measurement captures only one sales channel.
  4. Do the market selection. For US campaigns, DMAs are a common starting point, although states, cities, postal-code clusters, and other non-overlapping regions can also work. The goal is to create treatment and control groups with similar historical behavior before the test starts. Use historical sales or conversion data to check how closely the markets move together over time. Also check for spillover. Commuting, cross-border shopping, travel, or overlapping media coverage can expose control markets to treatment activity and weaken the comparison.
  5. Run a power analysis. Statistical power is the chance your test detects a real effect if one exists. Free tools like Meta's GeoLift and Google's GeoX, simulate your test on past data and tell you the smallest lift it could reliably see. That number is the minimum detectable lift, or minimum detectable effect (MDE).

If your expected lift is smaller than the MDE, add markets, run longer, or change spend more sharply. Recast recommends at least 80% power before launch.

The three test types from step 2:

Test typeWhat you doQuestion it answersExample
HoldbackLaunch a new channel in test markets onlyIs this new channel worth adding?Run CTV in 20% of markets before a national rollout
Go-darkTurn off existing spend in test marketsWhat do we lose if we cut this?Pause branded search in a set of markets
Heavy-upAdd spend in test marketsWhat do we gain if we spend more?Double YouTube budget in selected markets

Multi-cell experiments run more than one of these at once, such as two spend levels against one control. They need even more markets, so start with a single cell.

Read next: CTV Versus Social: When and How to Shift Your Ad Budget

Phase 2: Run the Test (4 to 6 weeks)

  1. Freeze everything else. No new promos, pricing changes, creative swaps, or bid strategy changes in test or control markets. What goes wrong: a mid-test discount in three markets makes the whole result unreadable.
  2. Set the geographic targeting in every platform. Exclude control markets in Meta Ads Manager, Google Ads, and TikTok location settings, and switch off audience features that ignore location. What goes wrong: national TV, PR, or always-on campaigns that you can't geo-target still reach the control markets, so pause them or spread them evenly across both groups.
  3. Log anything unplanned. Stock-outs, local news, weather, or a competitor's launch in one market all need a note so the analysis can account for them.
  4. Don't call the result early. What goes wrong: checking daily and stopping on a good week inflates false wins.

Phase 3: Measure the Result

  1. Hold a cooldown. Keep reading sales for one to two weeks after the spend change ends, so delayed purchases count.
  2. Run a placebo check. Re-run the analysis on a pre-period window where nothing changed. It should show no lift. In one published worked example, a placebo window showed a +1.3% lift that wasn't significant, which is what a clean setup looks like.
  3. Pull the result from your tool and read it with the framework in the results section below.

Before launch, run this checklist:

  • One decision written down, with the action you'll take for each outcome
  • Test type chosen and spend change sized
  • KPI agreed, with a data source for every sales channel
  • Markets matched on sales history, with buffer markets excluded
  • Power analysis passed at 80% or higher
  • Freeze list signed off by media, promotions, and creative teams

The setup work in Phase 2 is where most tests quietly break, because it touches every platform at once. Our paid media team runs the geo exclusions and the freeze checklist as part of the media plan.

How Long Should a Geo Lift Test Run, and How Many Markets Do You Need?

Plan for a test of 4 to 6 weeks across 10 to 20+ markets, with at least 4 to 5 times the test length in clean sales history. Then let the power analysis adjust those numbers for your data.

The guidance below compares recommended test duration, geographic coverage, and historical data requirements across major first-party and open-source geo-testing frameworks.

SourceMinimum test lengthTest regions and control regionsSales history needed
Meta GeoLift best practices15 days (daily data) or 4 to 6 weeks (weekly data), and at least one purchase cycle20+ geographic units25+ pre-test periods, 4 to 5x the test length, ideally 52 weeks
Google Meridian GeoXNo fixed minimum; moving from 4 to 6 or 8 weeks can lower the MDE when daily data is volatile10 geos minimum, 50 to 100+ for best resultsNot specified
Google Ads geo conversion liftSet per studyGoogle Marketing AreasRated by a feasibility score; Google advises against running on "Low"

Two patterns stand out. More markets and longer tests both shrink the smallest lift you can detect. And daily data beats weekly data, because it gives the model more points to learn from in the same calendar time.

How Much Does a Geo Lift Test Cost?

There is no fixed price for a geo lift test. The cost depends on the test design, market coverage, expected lift, sales volatility, and the amount of media change needed to create a detectable signal.

For some brands, the main cost is revenue put at risk in a holdout or go-dark test. For others, it is the extra media spend required for a heavy-up test.

Cost lineWhat it isHow to estimate
Opportunity cost (go-dark or holdback)Sales the paused channel would have driven in the test marketsTest markets' share of sales × expected incremental rate × test weeks
Added spend (heavy-up)Extra budget in test marketsThe test investment levels your power analysis recommends
Local media premiumNarrow geo targeting can raise CPMs and frequencyCompare CPMs in a short pre-test flight
Tool or vendor feeFree open-source tools, or a paid platformOpen-source: $0 plus analyst time; vendors price per test or by subscription
Analyst and team timeDesign, setup, monitoring, readoutPlan two to four weeks of part-time work across measurement and media

A quick worked estimate: if your test markets are 15% of sales, you expect the channel to drive 10% of sales there, and the test runs six weeks, the opportunity cost is about 1.5% of six weeks' revenue in those markets. That's small enough for most finance teams to approve once they see what the answer is worth.

Small tests can be cheap. The worked example cited earlier treated 3 cities for 21 days with 10 control cities and needed about €3,038 of spend to detect a lift of roughly 5%. The trade-off is precision: small tests can only see big effects.

Keep the test inside your planned channel budget where possible. A go-dark test saves money during the test, which usually makes approval easier.

How to Read Geo Lift Test Results

Read a geo lift result by looking at the range first and the headline number second.

A reported 12% lift with a range of 2% to 22% tells you something different from a 12% lift with a range of -5% to 29%. In the first case, the estimated effect stays positive across the range. In the second, the data cannot rule out little or no positive effect.

This distinction matters more than the headline number alone.

Calculate Lift and iROAS

Two numbers appear in most geo lift readouts:

Lift %

(Actual result - Predicted result without the media change) ÷ Predicted result × 100

Suppose the synthetic control predicts $1 million in revenue, but the test markets generate $1.1 million.

The estimated incremental revenue is $100,000, or a 10% lift.

iROAS

Incremental revenue ÷ Incremental media spend

If the test generated $100,000 in incremental revenue from an additional $50,000 of media spend, the iROAS is 2.0x.

This means each additional dollar of tested media generated an estimated $2 in revenue that would not otherwise have occurred.

Four numbers matter:

  • Incremental lift: The percentage increase in sales in the test group above what the synthetic control predicted.
  • Incremental conversions or revenue: The same lift expressed as orders or dollars.
  • iROAS: Incremental revenue divided by the spend you changed. An iROAS of 2.0 means each dollar returned two dollars that wouldn't have happened otherwise.
  • The range: Usually confidence intervals (or credible intervals in Bayesian tools), showing where the true effect likely sits.

An inconclusive result isn't proof of zero. It means the true effect is smaller than the test's MDE, the minimum detectable lift set in the power analysis. Treat it as "below X%," which is still useful if X is below your breakeven.

Decide What to Do With Each Result

What you seeWhat it meansWhat to do next
iROAS above target, range clear of zeroThe channel drives profitable incremental salesScale in steps and retest at the higher spend
Positive lift, iROAS below breakevenThe channel works, just not profitably at this spendCut spend, restructure campaigns, or change creative
Range crosses zeroThe test couldn't separate the effect from noiseRerun with more markets, longer duration, or a bigger spend change
Negative liftUsually a setup problem before a real effectCheck logged anomalies and spillover first, then rerun

Real results show how far this can move a budget. In Google's 2011 geo experiment paper we shared above, an advertiser's reported cost per click was $2.40 while the true cost per incremental click was $3.00.

Your own result can land far from the median. Polish fashion brand Reserved ran a geo lift test, increasing TikTok spend in selected regions and comparing them with a synthetic control. TikTok drove a 4.5% online revenue lift at 4.1x iROAS, and a 3.1% omnichannel lift at 5.6x iROAS once store sales were counted.

Pro tip: Write the decision rule before the test launches, such as "scale if iROAS is above 1.8 and the range stays above 1.0." Teams that set the bar after seeing the number tend to move it. Then translate iROAS into incremental cost per acquisition so it sits next to the rest of your customer acquisition metrics.

Can You Geo Test Influencer and Creator Spend?

Creator programs can be geo-tested when the creator content runs as paid media, because paid placements can be targeted by region and organic posts can't. That's the main constraint, and it's why most guides skip creators entirely.

An organic creator post reaches followers wherever they live. A creator in Austin with a national following leaks into every control market, so organic posting alone can't be read with a geo test.

Two designs work:

  • Paid creator ads with a regional holdout. Run creator content through whitelisted ads, Spark Ads, or partnership ads, and exclude the control markets. Running creator content as dark posts makes this practical, since each ad can carry its own location targeting.
  • Regional creator seeding. Recruit local creators with mostly local audiences in a set of test markets, then read retail or Amazon sales against matched market controls. This needs creators whose audiences you can confirm by region, using the regional breakdowns in each platform's creator analytics.

The first design is cleaner and easier to scale, because the same location exclusions from Phase 2 apply.

The second design suits brands with physical retail. We've used city-specific casting before: for New Balance, we matched athlete creators to its California stores, with 7 creators producing 50+ assets for local in-store promotion. A geo test layers measurement on top of that kind of local program.

Four New Balance athlete creators photographed outdoors in New Balance activewear: on a soccer pitch, on a bridge, by a river and on city steps.

Read next: 6-Step Plan to Optimize Paid Media Spend Across Multiple Channels

How to Report Geo Lift Results to Finance and Update Your Budget

Report a geo lift result to finance as one page with five parts:

  1. The decision tested. What budget question was the experiment designed to answer?
  2. The test design. Which markets changed, which stayed as controls, and how long did the test run?
  3. The result. Show incremental revenue, lift, iROAS, and the uncertainty range.
  4. The decision. State what will change because of the result.
  5. The budget impact. Show what that decision means for the next planning period.

Finance does not need every diagnostic chart. It needs enough evidence to understand the decision and the financial consequence.

That trust is already there. A Haus survey reported by EMARKETER found that 60% of senior decision-makers trust independent incrementality testing most, ahead of MMM at 40% and in-platform reporting at 37%.

Use the Result in Three Places

1. Channel budgeting

Compare the measured iROAS with the channel’s breakeven threshold and the returns available elsewhere in the media mix.

If the test produced profitable incremental revenue, that supports further investment. If the measured return fell below the required economics, reduce or restructure spend before adding more budget.

Avoid treating one test result as a permanent channel score. Incrementality can change as spend, audience saturation, creative, competition, and market conditions change.

2. Attribution and forecasting

Geo testing can show how far platform-reported performance differs from measured incremental performance.

For example, you can calculate:

Incremental conversions ÷ Platform-reported conversions

This gives you an incrementality ratio for the specific test conditions.

Use that ratio as a planning input, not a universal correction factor. Re-test it after material changes in spend, targeting, campaign structure, or market conditions.

This matters more as platform attribution becomes less complete. Meta, for example, removed several view-through attribution windows from Ads Insights reporting in January 2026, including 7-day and 28-day view metrics.

3. Media Mix Modeling

Experiment results can also inform an MMM model.

Bayesian MMM frameworks such as Google Meridian can incorporate experiment evidence as priors, helping the model stay closer to observed causal results rather than relying only on historical correlations.

Turn the Result Into the Next Budget Question

A geo lift test should close one decision and sharpen the next one.

If Meta prospecting cleared the required iROAS threshold, the next question might be how far spend can increase before incremental returns weaken. If branded search showed little incremental value, the next test might examine a smaller budget or a different campaign structure.

This creates a measurement cycle:

Test → decide → reallocate → retest.

Build a Measurement Plan You Can Defend with Fieldtrip

Geo lift testing gives you a number your finance team can trust. It works best inside a wider system: creative testing to find the ads worth scaling, media buying to distribute them, and measurement to prove which spend caused the sales.

That's how we run accounts at Fieldtrip. Our measurement analysts design the test and read the result, our media team handles the targeting and the freeze, and creator content feeds the channels being tested. Your first-party data stays at the center, so the answers hold up as platforms change their rules.

If you're weighing a scale or cut decision on a channel, contact us to plan your first geo test. We'll review your sales data by region, check whether a geo test can detect the lift you expect, and sketch a first test design.

FAQs

Can a brand selling in only a few regions still run a geo lift test?

Yes, with a smaller design. Google's time-based regression method was built for tests with very few geos, including one test market against one control. Expect to detect only large effects, and use finer units such as cities or ZIP-based markets to get more comparisons.

Does a geo lift test work when most sales happen on Amazon or in retail stores?

It does, if you can get those sales by region. Amazon sellers can usually pull orders by ZIP or state, and retailers share point-of-sale data by store or market, though often with a lag of a week or more. Build the lag into the cooldown and confirm the data feed before you launch.

How often should you rerun a geo test on the same channel?

Rerun when something material changes: spend rises by roughly 30% or more, the creative strategy shifts, or the market changes, for example during peak season. Many teams also retest each major channel once or twice a year, because incrementality drifts as audiences saturate.

Can a geo lift test measure brand awareness, or only sales?

It can measure any outcome you can collect by region, including branded search volume, site visits, or survey-based awareness. Sales remain the strongest KPI because they connect straight to iROAS. Upper-funnel metrics are useful as a secondary read when the purchase cycle is too long for the test window.

Can you run geo lift tests on two channels at the same time?

Yes, if the two tests use separate sets of markets or a multi-cell design that accounts for both. Running both tests on the same markets mixes the effects together and makes neither result readable. Most teams start with one channel at a time until the process runs smoothly.

David Morneau
David Morneau
Co-founder & CEO, inBeat Agency · CEO, Fieldtrip

David Morneau is the co-founder and CEO of inBeat Agency and CEO of Fieldtrip, the agency network that includes inBeat. Based in Montreal, Canada, he is a law graduate turned serial entrepreneur whose work spans paid media, performance creative, and search engine optimization (SEO).

View LinkedIn Profile →
Summarize this article with AI
ChatGPT
Perplexity
Claude
Grok
Table of contents
Add Fieldtrip as a preferred source on Google
Let us know what you’re working on.
We’re open to the right projects.
Let's Talk
A woman with red hair is laying on a bed.A woman wearing a white shirt with a dinosaur on it is sitting on a basketball hoop.Man sitting on a couch in front of a laptopTwo men are posing for a picture with a basketball.Women scrolling through phoneA woman in a purple shirt and black pants is posing for a picture.A woman wearing glasses and a green shirt is sitting in front of a laptop.A person is holding a piece of food in front of a plate of food.A woman with a nose piercing wearing a white shirt.A woman with curly hair and a white shirt.A person wearing a white shirt and a red shirt is using a laptop.A man and woman are smiling and hugging each other.A woman is sitting in front of a computer screen.A blue computer mouse on a mouse pad with a picture of a beach.A man drinking a beverage from a white bottle.A man holding a football and drinking a green smoothie.A man in an orange jacket is running.A woman is holding a bag full of clothes and is smiling.A man playing a trumpet on a concrete wall.A man with a beard and a Thai Larose shirt.A box of Dr. Squatch fresh falls men's natural soap.A man wearing a yellow jacket and black pants standing on a snowy hill.A person holding a cell phone with the word "Hoppers" on the screen.