What do the IAB incrementality guidelines actually require?
The Guidelines for Incremental Measurement in Commerce Media, published November 3, 2025 by IAB and IAB Europe through the IAB Commerce Board and its Incrementality Task Force, is a methodology framework.
Separately, IAB Europe released its Commerce (Incl. Retail) Media Measurement Standards V2.1. The latter is certifiable through their Retail Media Certification Programme with independent audits.
In other words, the IAB’s incrementality guidelines are exactly that – guidelines without any certification or audit attached. The Commerce (Incl. Retail) Media Measurement Standards are certifiable.
That matters because the two different documents are getting conflated.
The guidelines define incrementality as the causal impact of marketing, meaning the outcomes a campaign drove versus what would have happened without it. It draws a very clear distinction here: attribution and ROAS tell you what happened, not whether marketing caused it.
They also note a critical misconception: incrementality isn't static. The document reinforces a point we’ve made consistently, because competitive activity and consumer behavior are dynamically changing and interacting with media efficacy incrementality will change based on these conditions of commerce.
What makes a method causal?
Three requirements, and a method has to clear all three. 1) A credible counterfactual, meaning a valid "what if not" scenario, or a defined intervention separating treated from untreated; the guidelines say the rigor of this design determines the causal strength of everything downstream. 2) Control of bias, which they're realistic about, since you can't remove it all, only reduce it to where results stay actionable. And 3) separation of signal from noise, because a lift estimate sitting inside the noise floor is unusable no matter how good the design was. For that last one they want confidence intervals excluding zero, plus falsification and sensitivity tests, and they make the point that repeated measurement over time is how you tell a real effect from a lucky one.
The cornerstone of the guidelines is a hierarchy of measurement.
The four families, ranked
| Family | Examples | Causal Strength |
|---|---|---|
| Experiment-based | RCTs, holdouts, ghost ads, matched markets | Strong |
| Model-based counterfactual | Synthetic control, ML propensity models | Strong to moderate |
| Econometric | MMM, time-series regression | Moderate to weak |
| Hybrid proxies | New-to-brand percentage, baseline vs exposed, platform-reported incrementality, simple MTA | Weak |
That bottom rung of the ladder are platform-reported proxies, including metrics like new-to-brand percentage. These are the most ubiquitous metrics available today but also the weakest at drawing causal inferences on the effect of advertising. The reasoning is that those methods compare performance on related indicators or lean heavily on assumptions instead of independently estimating a counterfactual or directly measuring a control group.
Moving up the ladder are the econometric techniques like those that underpin marketing mix models (MMM). There are more holistic than proxy-based metrics insofar as they capture long-term aggregate effects and attempt to control for interaction with non-marketing factors such as seasonality. However, the Guidelines correctly point out limiting factors like multicollinearity and lack of granular predictive power hold back this method.
The next two types generate counterfactuals but do so through different mechanisms. Model-based counterfactuals attempt to predict or estimate an unobserved counterfactual or synthetically created one. Experiment-based explicitly establishes a measurable counterfactual through experimental design, holding back treatment from a group and comparing their behavior to those that were treated.
Strength of method vs. use cases
The guidelines go to great lengths to highlight that strength of approach is not the only consideration. One must also consider the business objective of measurement and the constraints that places on methods. While they highlight a number of different use cases, in general business utility or actionability cuts in the opposite direction of measurement strength.
Designed experiments are strong on causal inference but narrow in reach, usually one platform at a time. MMM is weaker on causal inference but broader on scope, covering every channel plus non-media factors.
There isn’t a single method that should be applied universality rather the method must be balanced against the requirements of the use case.
What to check on your own program?
Sort your available measurement into the four families and map those against the types of decision you are using each for. Are you using hybrid proxies to make large budget allocation decisions across channels? That’s a dangerous mismatch of method and use case.
Are you running designed experiments but trying to apply them to optimize always on campaigns? A different method may be required.
Use this as a starting point for conversation both internally with analytics and media planning teams, with agencies, with retail media networks and any 3rd party measurement partners. How can each of these teams better align their individual use cases to the appropriate measurement.
Our POV
We think the guidelines and framework presented within it are a key foundational step towards the adoption of more robust measurement within commerce media, and we applaud the IAB and the contributing parties for the development of these guidelines.
In our view, the crux of the problem presented by it is this tension between balancing managerial utility and measurement strength. Our own approach to resolving this is an ensemble model. By blending the strengths of multiple techniques some of these tradeoffs can become ANDs, not ORs. Particularly today, where designed experiments can only be run by the retailers/RMNs themselves, limiting the availability of this technique to be deployed at scale, a multitude of techniques is required to achieve both robust measurement and scale. There is already a general shift towards “triangulation” in the broader marketing effectiveness and measurement space and believe it is essential here in commerce media as well.
The one area we’d highlight further is the need to bring a greater context into the conditions of commerce into all of these techniques. Factors outside of media such as competitive shelf presence, pricing and ratings/reviews greatly affect consumer behavior both online and in-store. These factors don’t just need to be controlled for in experimental design but incorporated foundational into an understanding of media - how does running a promotion affect consumer’s responsivity to advertising? How does improving listing quality improve the efficiency of advertising? This unlocks an understanding of how to effectively leverage organic and paid levers together to generate incremental growth.
The future of commerce and media requires a unified causal intelligence layer to power these use cases, particularly as these decisions become increasingly mediated by agentic AI.