B
Brenden Delarua
Guest
Most marketing teams measure the same channel several ways and get several different answers. Last-click attribution says one thing, the platform dashboard says another, the media mix model says a third.
At Stella, we reconcile them with something we call the Calibration Loop. Incrementality experiments are the strongest causal anchor we have, so we use them to inform a media mix model that fills in the periods and spend levels the experiments never directly see, and we keep attribution and post-purchase surveys running as coverage in between. We are not trying to pick the one true method. We are trying to let the strongest evidence keep the weaker evidence honest.
They rest on different assumptions, and most of what runs your budget is observational. Attribution and platform dashboards read patterns in the data as it comes. Targeting and delivery make the people who saw your ad different from the people who didn't, so a conversion that happens after an ad isn't automatically a conversion the ad caused. A media mix model can be causal too, but only under assumptions you have to state and defend out loud. Randomized experiments carry the lightest assumptions of the three, which is why we lean on them hardest.
The gap between observational and causal is not small. In a well-known set of 15 Facebook field experiments, observational estimates of purchase lift were off by at least a factor of three in half the studies (Gordon et al., 2019, Marketing Science). The scary part isn't the size of the miss. It's that without a causal design you have no idea which way the bias runs. Your reporting could be overstating a channel by 200%, or understating it, and you can't tell which from the dashboard.
It's a sequential learning system, not one tool. The experiments tell us what we can know directly. The model estimates what happens in between. The next experiment goes wherever that estimate is weakest, or wherever the decision riding on it matters most. Then the whole thing runs again.
Using experiments to calibrate a media mix model isn't something we invented. Meta's Robyn and Google's Meridian both support it, and any serious measurement team is doing some version of it already. The part we actually care about is the loop around it: which test to run next, how you read it when two methods disagree, and what sends you back to re-test.
The observable isn't the incremental. And the incremental isn't necessarily the marginal. Keeping both of those in your head at once is the entire reason the loop exists.
A randomized holdout gives you the cleanest causal read you can get in practice. Withhold exposure from a comparable group, measure the difference, and the number you get doesn't need a model to exist.
What it does need is context, and this is where the industry oversells it. A single experiment is causal for one treatment, one audience, one geography, one spend level, one stretch of time, one creative mix, with real sampling error on top. Push that result into a model that has to cover other periods and other budgets, and you are adding uncertainty, not removing it. The experiment is the thing the model gets held against. It doesn't make the rest of the model automatically correct.
Holdouts aren't the only way to get at a counterfactual, either. Synthetic controls, difference-in-differences, and the other quasi-experimental designs all estimate it under their own assumptions. We reach for randomized holdouts when they're feasible and properly powered, and use the alternatives when they aren't.
They're coverage. You can't run a holdout on every channel every week, so attribution and post-purchase surveys keep eyes on the map between experiments. Both come with baggage. Attribution credits whoever showed up before the sale, whether the ad moved them or not. Surveys have recall bias, selection bias, people telling you what they think you want to hear. Neither one gives you a clean causal number, and we don't pretend they do.
Surveys earn their spot because they catch exposure nothing else can see. Word of mouth. A podcast someone half-remembers. Channels that never leave a clickable trail. That's coverage the models are blind to. It still isn't a causal read.
The interesting case is when a survey and an experiment flat-out disagree. Say the survey gives podcasts credit for a third of orders and the holdout shows barely any lift. The experiment is the one estimating cause and effect. The survey is telling you there's a pathway worth digging into. You treat that gap as a lead to chase. You don't let the survey overrule the test.
Knowing your current spend was incremental still doesn't tell you what the next dollar does, and the next dollar is what a CFO is really asking about at budget time. So the model estimates response curves and the uncertainty around them, and that's where the marginal read comes from. What we don't do is spend the money. We measure, hand back validated incremental revenue and the shape of the curve, and the buying team or the platform takes it from there.
Keeping measurement separate from media buying kills one obvious incentive to fudge the numbers. It does not make the model right. Plenty of independent vendors build bad models. Independence buys you cleaner incentives and nothing else, so don't let anyone sell it to you as a guarantee.
In the Q2 2026 YouTube Ads Report, which we co-produced, 92 holdout studies covering $2.3M in test spend came back with a median 2.01x incremental ROAS on YouTube and Demand Gen. The median is the least useful thing in that sentence, though. Incrementality swings hard by brand, by spend level, by creative, so one median paves over most of what's actually going on. The spread is the story.
Our own book says the same. Here's the caveat that has to come first: these are brands that chose to run a test. Read it as evidence, not as a market average. Across 225 incrementality tests on DTC brands, the median incremental ROAS is 2.31x, and the middle half of results land between 1.36x and 3.24x. That range is the whole point. You measure incrementality per case. You don't get to assume it going in.
Don't reflexively start with your biggest channel. Start where a better number would actually change a decision, and where you can build a test with enough power to trust the answer. A giant channel you can't cleanly randomize, or one that's already well identified, makes a worse first test than a mid-sized one where the result would move budget.
Then you've got a choice about which number you defend next quarter. The one the platform handed you, or the one that lived through a holdout and showed up with an honest error bar. I know which one I'd want in the room.
Photo by H&CO on Unsplash
At Stella, we reconcile them with something we call the Calibration Loop. Incrementality experiments are the strongest causal anchor we have, so we use them to inform a media mix model that fills in the periods and spend levels the experiments never directly see, and we keep attribution and post-purchase surveys running as coverage in between. We are not trying to pick the one true method. We are trying to let the strongest evidence keep the weaker evidence honest.
Why the numbers disagree.
They rest on different assumptions, and most of what runs your budget is observational. Attribution and platform dashboards read patterns in the data as it comes. Targeting and delivery make the people who saw your ad different from the people who didn't, so a conversion that happens after an ad isn't automatically a conversion the ad caused. A media mix model can be causal too, but only under assumptions you have to state and defend out loud. Randomized experiments carry the lightest assumptions of the three, which is why we lean on them hardest.
The gap between observational and causal is not small. In a well-known set of 15 Facebook field experiments, observational estimates of purchase lift were off by at least a factor of three in half the studies (Gordon et al., 2019, Marketing Science). The scary part isn't the size of the miss. It's that without a causal design you have no idea which way the bias runs. Your reporting could be overstating a channel by 200%, or understating it, and you can't tell which from the dashboard.
What the loop actually is
It's a sequential learning system, not one tool. The experiments tell us what we can know directly. The model estimates what happens in between. The next experiment goes wherever that estimate is weakest, or wherever the decision riding on it matters most. Then the whole thing runs again.
Using experiments to calibrate a media mix model isn't something we invented. Meta's Robyn and Google's Meridian both support it, and any serious measurement team is doing some version of it already. The part we actually care about is the loop around it: which test to run next, how you read it when two methods disagree, and what sends you back to re-test.
The observable isn't the incremental. And the incremental isn't necessarily the marginal. Keeping both of those in your head at once is the entire reason the loop exists.
Why experiments come first
A randomized holdout gives you the cleanest causal read you can get in practice. Withhold exposure from a comparable group, measure the difference, and the number you get doesn't need a model to exist.
What it does need is context, and this is where the industry oversells it. A single experiment is causal for one treatment, one audience, one geography, one spend level, one stretch of time, one creative mix, with real sampling error on top. Push that result into a model that has to cover other periods and other budgets, and you are adding uncertainty, not removing it. The experiment is the thing the model gets held against. It doesn't make the rest of the model automatically correct.
Holdouts aren't the only way to get at a counterfactual, either. Synthetic controls, difference-in-differences, and the other quasi-experimental designs all estimate it under their own assumptions. We reach for randomized holdouts when they're feasible and properly powered, and use the alternatives when they aren't.
What attribution and surveys are for.
They're coverage. You can't run a holdout on every channel every week, so attribution and post-purchase surveys keep eyes on the map between experiments. Both come with baggage. Attribution credits whoever showed up before the sale, whether the ad moved them or not. Surveys have recall bias, selection bias, people telling you what they think you want to hear. Neither one gives you a clean causal number, and we don't pretend they do.
Surveys earn their spot because they catch exposure nothing else can see. Word of mouth. A podcast someone half-remembers. Channels that never leave a clickable trail. That's coverage the models are blind to. It still isn't a causal read.
The interesting case is when a survey and an experiment flat-out disagree. Say the survey gives podcasts credit for a third of orders and the holdout shows barely any lift. The experiment is the one estimating cause and effect. The survey is telling you there's a pathway worth digging into. You treat that gap as a lead to chase. You don't let the survey overrule the test.
The Next Dollar
Knowing your current spend was incremental still doesn't tell you what the next dollar does, and the next dollar is what a CFO is really asking about at budget time. So the model estimates response curves and the uncertainty around them, and that's where the marginal read comes from. What we don't do is spend the money. We measure, hand back validated incremental revenue and the shape of the curve, and the buying team or the platform takes it from there.
Keeping measurement separate from media buying kills one obvious incentive to fudge the numbers. It does not make the model right. Plenty of independent vendors build bad models. Independence buys you cleaner incentives and nothing else, so don't let anyone sell it to you as a guarantee.
What it looks like with real numbers
In the Q2 2026 YouTube Ads Report, which we co-produced, 92 holdout studies covering $2.3M in test spend came back with a median 2.01x incremental ROAS on YouTube and Demand Gen. The median is the least useful thing in that sentence, though. Incrementality swings hard by brand, by spend level, by creative, so one median paves over most of what's actually going on. The spread is the story.
Our own book says the same. Here's the caveat that has to come first: these are brands that chose to run a test. Read it as evidence, not as a market average. Across 225 incrementality tests on DTC brands, the median incremental ROAS is 2.31x, and the middle half of results land between 1.36x and 3.24x. That range is the whole point. You measure incrementality per case. You don't get to assume it going in.
Where to start.
Don't reflexively start with your biggest channel. Start where a better number would actually change a decision, and where you can build a test with enough power to trust the answer. A giant channel you can't cleanly randomize, or one that's already well identified, makes a worse first test than a mid-sized one where the result would move budget.
Then you've got a choice about which number you defend next quarter. The one the platform handed you, or the one that lived through a holdout and showed up with an honest error bar. I know which one I'd want in the room.
Photo by H&CO on Unsplash