Experiments are a moving target in retail media
Experiments are often considered the gold standard of measurement, but there are two assumptions underneath that get overlooked: 1) that a cleanly designed experiment can be run, and 2) that the results of the test can be generalized to apply to future decisions.
Today designed experiments across the major RMNs are only available as a managed service offering – meaning advertisers can’t design and run the experiments themselves at scale; they are wholly reliant on the RMN to design and run them. The net result is that actual designed experiments, whether randomized controlled or geo-based, are often few and far between.
This has led to a lot of advertisers trying to creatively run simple pre/post designs, where there is no parallel unexposed group, just a pre-period prior to treatment.
But trying to find or design a clean experiment like that is nearly impossible in the constantly changing dynamics of Amazon’s marketplace. Amazon’s advertising FAQ says a listing has to be eligible for the Featured Offer for its ads to run, and that a product has to be in stock and priced competitively to hold the Featured Offer. If it isn't, the ad "will not display." Let’s say a competitor undercuts you by 10% for a week, a replenishment order lands late, or somebody splits a variation and four thousand reviews scatter across five child ASINs. Conversions move, and on Amazon delivery moves with it, because the ads stop the moment the Featured Offer goes, and with it your well-designed experiment.
The error runs the other way too, and that version is dangerous, because nobody audits a result they like. If you happen to run a test during a window when a pricing correction lands or a review push is finally cleared, if those aren’t eveningly distributed across your test design, it can inflate results.
This makes it tempting to think you need to strip the effects of changing shelf conditions out of your design entirely, so the media effect shows up clean underneath, but that’s not quite right either.
Some shelf movement really is noise. A competitor's timely promo or a listing suppressed in four states has nothing to do with your media plan, yet it shows in your media’s performance. But a lot of shelf movement is your media. Ads drive velocity, velocity moves organic rank, better rank drives sales and those sales pull in reviews. That stronger rank and a fuller review count will generate additional impact stretching outside of the test. A test that stripped out every shelf effect would undercount the exact thing you're trying to size.
The platforms aren't blind to any of this, incidentally. Pacvue offers rules that automatically adjust advertising when inventory drops or Buy Box ownership changes, and Skai has their own version. That stops you pouring budget into a listing that can't convert, but it does nothing for your read afterward. A rule that pauses spend when a star rating slips ends up cutting your media for a reason that already predicts weaker conversion, so both curves fall together and neither one caused the other.
So how should it be done?
The obvious answer is self-service experimentation allowing for randomized controlled triasl to be run at scale.
But even if tomorrow the RMNs unlocked this, running a clean experiment is only half the battle.
If those experiments can’t be generalized and applied to future decisions, their utility is still limited. If the experiment was run while the listing had low organic rank and limited reviews, can it be applied to a situation that has different conditions? What happens if pricing dynamics are different now then they were during the experiment? Are the results still applicable to current decisions?
These marketplaces are constantly changing, with search dynamics, pricing, promotions, and ratings and reviews all in flux and all affecting advertisers’ ability to influence consumer behavior. And that makes the practice of “weighting” last-touch ROAS with experimental results extremely problematic in commerce, because it assumes these conditions have not changed since the experiment was run.
If an experiment was run while you had competitive pricing in the category – even if the test stripped out the impact of price to isolate media’s effect – if two weeks later competitors drop their price and you are no longer competitive, that lift estimate is no longer a useful estimate of media’s impact today. No amount of media will overcome that pricing disadvantage.
This highlights an important methodological consideration in commerce: experiments are an input not the final answer for decisioning. That’s why experiments are just one part of the approach we take at Incremental. Designed experiments alongside quasi-experimental approaches like difference-in-differences provide a robust historical causal foundation. Those are fed into a SKU-level econometric model that estimates how sales and media respond to the changing shelf conditions. This allows the model’s understanding of media incrementality to respond in real-time to changes in price, inventory, search rank, promotion, and competitive activity, so shelf conditions enter the model as influencing features rather than contaminating factors.
The payoff is a far more robust real-time signal to use for day-to-day decisioning like in-flight campaign optimization. It keeps pace with your media and can be used while budgets are still moving, instead of arriving later, after the dust has settled.
It also makes comparisons of performance across retailers and across the funnel far more equitable. The differences in these commerce dynamics between retailers aren’t penalizing campaign performance, allowing you to compare the performance of a campaign on one RMN where you have strong organic ranking, to one where you have weak organic visibility.
The fact is, the conditions of commerce change. But that isn't noise to be engineered out, it's part of the foundational factors influencing consumer behavior and an essential input to media decisioning.