Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

For years, observational studies found that vaccinated seniors were about half as likely to die over the winter as unvaccinated ones. Taken at face value, that would make the flu shot one of the most powerful life-extending interventions in medicine, cutting deaths from every cause in half. It is not. The giveaway came from looking where the vaccine could not possibly help. When researchers split the deaths by timing, the largest mortality advantage appeared in the weeks before influenza was even circulating. A flu shot cannot prevent death from a virus that has not yet arrived. The apparent effect was not protection; it was that the people who show up to get vaccinated are healthier to begin with. The comparison was confounded, and the deaths in the wrong season exposed it.

That check is an instance of a general and underused idea: the negative control. A negative-control outcome is an outcome your exposure could not have caused, a place where you already know the true answer is zero. You run the same analysis on it. If it returns an effect anyway, your method is manufacturing effects, and the bias that produced the false one is almost certainly distorting your real estimate too.

The power of this comes from making the invisible visible. Residual confounding is normally impossible to see, because the whole problem is a variable you did not measure or did not think of. A negative control gives you one spot where any nonzero result has to be bias. It turns an untestable anxiety, maybe something is confounding this, into a falsifiable prediction: if my design is clean, the effect here should be zero. It is the negative control from the laboratory bench, the well that should show no reaction, imported into observational research.

It comes in two forms. A negative-control outcome is an outcome the exposure cannot affect but that is subject to the same selection, like death before flu season. A negative-control exposure is an exposure that cannot cause the outcome but carries the same confounding. The classic case asks whether a mother’s smoking in pregnancy harms the baby through the womb or through the hard circumstances that accompany smoking. Use the father’s smoking as the control: if it predicts the same harm, the pathway is shared family circumstance, not the exposure in the womb. Either way the rule is the same: the control must share the confounding of the real question while having no true causal link of its own.

Chosen well, negative controls can do more than flag a problem; newer methods use them to estimate the bias and subtract it. But the everyday value is diagnostic. It is a cheap and often decisive smell test that needs no new data, only a question you already have the data to answer.

This transfers straight into program evaluation, where negative controls usually hide in two familiar places. The first is the pre-period: did the improvement you are crediting to the program begin before the program did? An effect that predates its cause is a negative control failing in front of you. The second is the off-target outcome: did your job-training program also seem to improve something it cannot plausibly have touched? If the same participants look better on an outcome the training could not have moved, the flattering result on the outcome you care about is under suspicion too.

Two cautions keep this honest. A clean negative control is reassuring but not a clean bill of health, since it rules out the confounding that runs through that one channel, not every bias that might exist; and a control that does not actually share the confounding proves nothing, so choosing them well is a real craft. A failed negative control, however, is close to decisive. It is direct evidence, in your own data, that your method reports effects that are not there.

So here is my question.. Before you believe your own positive result, have you found the place where the answer must be zero, and checked that your analysis agrees?

Posted in

Leave a comment