A campaign can look effective in a dashboard and still be taking credit for demand that would have happened anyway. Incrementality testing gives a team a more useful answer. Instead of asking which touchpoint received credit, it asks what extra revenue, leads, or customers the activity actually created. That distinction matters whenever a budget decision is larger than a routine optimisation.

What incrementality testing measures

Incrementality testing is a controlled comparison. One group is exposed to a marketing activity, while a comparable group is not. If the exposed group produces a materially stronger outcome after the test, the difference is the evidence that the activity added something beyond normal demand.

For a retailer, the outcome may be new-customer revenue. For a lead-generation business, it may be qualified opportunities that reach a sales conversation. For an established brand, it could be subscriptions, store visits, or another meaningful action that connects to growth. The metric needs to be agreed before the test begins, otherwise the result can be bent to fit the story someone hoped to tell.

Google’s overview of incrementality testing frames the method around measuring the true contribution of marketing. That is the real value. A good test does not replace day-to-day reporting. It gives leadership a clearer basis for deciding where the next dollar should go.

Why attribution alone cannot settle the question

Attribution is useful. It helps a team understand the paths people take before they convert and can reveal which campaigns tend to appear near valuable actions. But attribution observes a journey after it has happened. It cannot fully answer whether a person would have purchased, submitted a form, or searched for the brand without the marketing activity.

This is especially important in channels that influence demand before a clear click occurs. Programmatic display, connected TV, video, retail media, and broader prospecting can all reach people who later convert through a branded search, direct visit, email, or another channel. A last-click report may understate their role. A platform report may overstate it. Neither view proves the incremental effect on its own.

The right question is not whether attribution is good or bad. It is whether it is strong enough for the decision in front of you. A small bid adjustment may only need disciplined campaign reporting. A major move across paid search and Shopping, programmatic advertising, and other media deserves evidence that is closer to cause and effect.

Use a test when the decision is consequential

Incrementality testing is not a ritual to run on every campaign. It works best when there is a meaningful choice to make and the organisation can realistically change course based on the result. Common questions include whether a prospecting campaign is generating new customers, whether a promotion is creating demand or merely pulling forward purchases, and whether an additional layer of media can scale efficiently.

It is also useful when a team is debating two plausible explanations for the same trend. Perhaps branded search rose after a streaming campaign. Perhaps it rose because of seasonality, a product launch, press coverage, or a competitor’s change in activity. A controlled comparison will not eliminate every uncertainty, but it is more rigorous than asking a dashboard to explain a market change it cannot observe.

Before planning the test, write one sentence that states the decision. For example: “Should we continue funding this connected TV audience after the initial launch?” Then define the business outcome that would support a yes, the amount of risk the company is willing to take, and the action each possible outcome would trigger. If no decision changes, do not run the test yet.

Start with a test and control group

The core design is simple: create a test group that can see the marketing activity and a control group that cannot. The hard part is making the two groups comparable. If one group already has stronger demand, a higher-income customer base, more stores, or a different buying cycle, the difference at the end of the test may reflect those conditions rather than the media.

Two equal groups of black markers divided by a physical barrier, with light falling on one group

A practical test plan should answer four questions before activity begins:

  1. What outcome matters? Choose one primary measure tied to the decision, such as qualified sales, new-customer orders, or revenue after returns. Supporting measures can help diagnose the result, but they should not replace the primary outcome.
  2. Who or where is eligible? Define the audience, geographic areas, stores, accounts, or customer segments that could reasonably be placed in test and control groups.
  3. What changes for the test group? Be specific about the campaign, budget, creative, targeting, offer, and timing. A vague “more media” test is difficult to interpret.
  4. What stays stable? Record promotions, price changes, distribution shifts, site releases, email sends, and other activity that could influence the outcome during the same period.

Google Ads describes Conversion Lift as comparing an exposed group with a comparable control group to estimate conversions driven by advertising. Its Conversion Lift guidance is a useful example of why the control group needs to be designed into the measurement, not added as an afterthought.

Choose the test shape that fits the business

There is no single “best” incrementality test. The right method depends on how customers buy, where the campaign can be controlled, and how much traffic or conversion volume is available. Choosing a smaller, clean test is usually better than attempting a broad study with no defensible comparison.

Audience holdouts work when a platform can randomly withhold activity from part of an eligible audience. They are often useful for digital campaigns, particularly when the platform has enough reach and the conversion event can be observed consistently. The test needs to avoid accidentally reaching the control group through another campaign with nearly identical targeting.

Geographic tests compare matched markets, regions, or store territories. They can work well for local retail, media that is bought by market, or businesses with offline sales. The challenge is selecting areas that behave similarly enough before the test. A high-growth market should not be compared with a market that is already flat or affected by a different local condition.

Abstract tile map with two distinct coloured market groups and routes

Time-based tests pause or change activity during selected periods and compare results with a baseline. They are easier to organise but more vulnerable to outside effects, such as a holiday, weather event, competitor promotion, or product change. They can still be useful for a narrow question, provided the team is candid about what the comparison cannot prove.

For ecommerce teams, a focused test can answer whether an incremental spend layer is adding profitable new demand rather than just collecting conversions from shoppers already on their way to purchase. That question pairs naturally with the commercial view on the Ecommerce PPC Agency page, where product priorities, paid media, conversion paths, and reporting are considered together.

Keep the test clean enough to trust

Most weak tests do not fail because the maths is difficult. They fail because too many things change at once. If the test group sees new creative, a new offer, a new landing page, and a new budget while the control group sees none of them, the result may be useful for the combined package. It cannot tell you which element did the work.

Protect the comparison. Keep pricing, inventory, site performance, fulfilment, sales follow-up, and major owned-media campaigns as consistent as possible. Monitor both groups during the test, but avoid changing the rules halfway through because an early result looks exciting or uncomfortable. Ending a test when a preferred answer appears is a reliable way to create confidence without insight.

It also helps to document exceptions as they happen. If a control market loses stock, a sales team changes its process, or a tracking event breaks, note the date and likely impact. A clear record lets the team decide whether the test should continue, be adjusted, or be treated as directional evidence rather than a final verdict.

Read lift in business terms

At the end of the test, start with the difference in the agreed primary outcome. Did the exposed group produce more qualified revenue, customers, or opportunities than the control group? Then ask whether the size of that lift justifies the spend, the operational effort, and the next level of investment.

A positive lift is not an instruction to scale without limits. Check whether the result is large enough to matter after cost, margin, returns, sales quality, and the expected drop in efficiency that can occur as spend grows. A campaign can create incremental conversions and still be the wrong place to put the next budget if the economics are weak.

A neutral or negative result is not wasted work either. It may show that the audience was already likely to convert, that the offer did not move the market, that the creative needs attention, or that the channel should play a different role. The point is to make the next decision better, not to defend the activity that was tested.

Two business leaders reviewing blank decision cards beside a small balance scale

Bring the finding into a short decision brief. State the original question, test design, primary outcome, result, known limitations, and recommendation. This is where incrementality testing becomes useful beyond a measurement team. Leaders can see what was learned, what remains uncertain, and why a budget should stay, move, or grow.

Common mistakes to avoid

Testing a metric that does not connect to the decision. Clicks, impressions, and platform conversions may help diagnose delivery, but they are rarely enough to prove business impact. Choose the closest reliable outcome that the organisation can act on.

Using unmatched groups. A test is only as credible as its comparison. Check historical performance, customer mix, geography, and seasonality before treating two groups as equivalent.

Measuring too soon. A short test window can miss the real purchase cycle, especially for considered purchases or longer sales processes. Agree on the observation window before launch.

Ignoring overlap. If the control group can still see the activity through another campaign, market, device, or channel, the difference between groups may be diluted. Map the full media plan before the test starts.

How Surge helps turn tests into better action

Surge connects the media plan, the measurement design, and the business decision in one working view. Our predictive data and analytics work helps teams define reliable outcomes and reporting, while our paid search and Shopping and programmatic advertising work creates the controlled activity a useful test requires.

That combination matters because the result of an incrementality test should change what happens next. Explore the case studies for examples of how sharper measurement and media decisions can improve the economics of growth, or talk with Surge about a campaign question that needs a more dependable answer.

A practical place to start

Choose one media decision that is large enough to matter and uncertain enough to deserve a better answer. Write the decision in plain language. Name the business outcome. Identify a test and control group that can be kept reasonably comparable. List the conditions that must stay stable, then agree on what result would justify a change in budget or approach.

That is enough to begin. The strongest incrementality programs are built through repeated, disciplined tests, not one oversized study that tries to answer every question at once. Start with the decision that has the most at stake, learn from it, and use that evidence to make the next move with more confidence.