Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

Knowing that a program worked is valuable. Knowing why it worked is more valuable still, because a mechanism you understand is one you can strengthen, cut, or carry to a new setting. So we naturally want to go further than the total effect and ask how much of it flowed through a particular pathway. Did the training raise earnings by building skills, by building confidence, or by signaling effort to employers? This is mediation analysis: splitting a total effect into the part that runs through a proposed mechanism and the part that does not. It is far more treacherous than its popularity suggests.

For decades the default has been the approach Reuben Baron and David Kenny laid out in 1986, one of the most cited recipes in social science. Run a few regressions. Estimate the effect of the treatment on the outcome, then on the mediator, then put the treatment and mediator into the outcome model together and watch the treatment coefficient. If it shrinks, you conclude the mediator carries part of the effect. The logic is intuitive, the steps are simple, and it is taught almost everywhere.

The problem hides in that last regression, where you control for the mediator. That gives a clean estimate of the mechanism only if you assume something strong: that nothing unmeasured causes both the mediator and the outcome. And here is the catch. Randomizing the treatment does not buy you that assumption. Randomization makes the treatment clean, so the treatment is unconfounded. But the mediator is never randomized. It is something that happened naturally after treatment, and the people who ended up with more of it, more confidence, say, may differ in unmeasured ways that also shape earnings. Those differences confound the mediator and the outcome even in a flawless experiment.

It gets sharper still, because the mediator is a post-treatment variable, and that is exactly the kind of variable an earlier post in this series warned against controlling for. If the mediator and the outcome share an unmeasured common cause, then adjusting for the mediator, precisely what the recipe instructs, opens a spurious path and introduces collider bias. Controlling for the wrong variable does not clean the estimate; it contaminates it. The standard recipe can invent a direct effect that is not real, or erase one that is, and a randomized treatment does nothing to prevent it.

Modern causal mediation analysis does not make these difficulties disappear, but it does make them honest. It defines the direct and indirect effects precisely, states the assumptions out loud, no unmeasured confounding of the treatment-outcome, treatment-mediator, or mediator-outcome relationships, and no mediator-outcome confounder that is itself affected by treatment, and it comes with sensitivity analysis to ask how badly a violation would have to bite before the conclusion flips. That is the discipline this series has urged elsewhere: you cannot make an untestable assumption true, but you can make it visible and probe how much weight it can bear.

The honest takeaway is a reversal of the usual instinct. A clean mediation result deserves more scrutiny than a clean estimate of the total effect, not less, because it rests on assumptions the study design cannot secure. Whether a program worked can sometimes be settled by a good experiment. How it worked almost never can be, at least not by the experiment alone.

For those of us who build logic models and theories of change, this matters, because those diagrams are full of mediation claims about which link is doing the work. Testing them quantitatively is worth doing, but it calls for naming the mediator-outcome confounders you fear, measuring the ones you can, and reporting how sensitive the mechanism claim is to the ones you cannot. A confident line like forty percent of the effect ran through this pathway should invite a hard look at what had to be assumed to say it.

So here is my question. When you claim that a program worked through a particular mechanism, do you hold that claim to a higher standard than the claim that it worked at all, or to a lower one?

Posted in

Leave a comment