“Does the program work?” sounds like the most basic question an evaluator can ask. It is also close to unanswerable as written, because the same program routinely succeeds in one place and fails in another, for reasons that a simple yes or no can never hold. The honest answer is almost always: it depends. Realist evaluation takes that “it depends” seriously and turns it into a method.
The starting point, set out by Ray Pawson and Nick Tilley in 1997, is that programs do not work the in the same way a drug is imagined to work. A program does not cause an outcome by itself. It offers resources, or opportunities, and what produces the result is how people respond to that offer. That response is the mechanism, and whether people respond one way or another depends on their context. The same job-training program can lift employment for one group and do nothing for another, not because the program changed, but because it set off different reasoning in a different setting.
This reframes the unit of analysis. Instead of asking for the average effect, realist evaluation asks for a configuration: in this context, what mechanism does the program trigger, and what outcome follows? The product of the evaluation is a set of these context-mechanism-outcome statements. In this context, for these people, this mechanism fired and produced this result; in that context, a different mechanism fired and produced something else. The aim is to explain the pattern, not just to detect it.
Contrast that with the familiar experimental question. A clean trial treats the program as a black box and estimates its average effect on the people in the study. That is valuable, but it tells you whether the box worked on average, in this sample, not why, for whom, or whether it will work anywhere else. A program with a healthy average effect can still be useless or harmful for a sizable subgroup, and a program judged a failure may have worked well wherever the context was right. The average can hide the mechanism completely.
The payoff for getting this right is transfer. When you know the mechanism and the conditions it needs, you have something you can carry to a new site and reason about in advance: this works when these conditions hold, so here is what to expect somewhere new. “The program works” does not travel. “This mechanism fires when these conditions are present” does. For a decision-maker weighing whether to scale or adopt, that is usually the more useful kind of knowledge.
None of this is free. Realist evaluation is demanding. It requires a theory of the mechanisms before you start, data on context and not merely on outcomes, and the discipline to test and refine your configurations rather than narrate them. It is less tidy than a single headline number, and it resists the clean summary that sponsors often want. It also carries a real risk: without specifying what would count as disconfirming, a set of context-mechanism-outcome stories can slide into unfalsifiable explanation. The rigor lives in stating the conjectures clearly and putting them to the data.
For those of us working in federal evaluation, this is not abstract. Programs run across wildly different sites, populations, and conditions, and a single effect averaged over all of them can be both technically correct and practically worthless. A program office deciding where to expand and where to pull back is rarely served by “it works.” It is served by knowing what works, for whom, under what conditions, and through what mechanism. That is the question worth the effort, even when it refuses to fit on one line.
So here is my question. When you report that a program works, do you also specify for whom and under what conditions, or does the average effect quietly stand in for the whole story?

Leave a comment