Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

Contrasting proxy metrics with real-world outcomes, including user satisfaction, revenue growth, and community impact

We rarely measure the thing we actually care about. Durable employment, real learning, long-term health, and safety are slow to arrive, costly to observe, and hard to pin on any one program. So we measure a proxy that is faster and cheaper, a job placement, a test score, a lab value, and we treat movement in the proxy as progress toward the goal. The proxy is usually chosen for a good reason: it tracks the outcome we want. The danger is that a program can move the proxy without moving the outcome, and sometimes move the proxy while pushing the outcome the wrong way.

The clearest evidence comes from medicine, where the stakes made the lesson unmissable. In the 1980s it was well established that irregular heartbeats after a heart attack predicted sudden cardiac death. Drugs that suppressed those arrhythmias were expected to save lives, and the expectation was so strong that many researchers considered withholding them in a trial unethical. The drugs were approved on the strength of the surrogate and taken by hundreds of thousands of people. When a placebo-controlled trial was finally run, they suppressed the arrhythmias exactly as designed and roughly doubled the death rate. The number improved; the patients did not.

The reason is subtle and general. A surrogate earns its place because it predicts the outcome; an arrhythmia predicts death. Predicting an outcome, however, is not the same as carrying a treatment’s effect on it. A drug can move the surrogate through one pathway and harm the outcome through another the surrogate never registers. Torcetrapib raised HDL, the so-called good cholesterol, and increased deaths, a real signal of risk and a false guide to treatment. This is a point an earlier post made about prediction and explanation: a variable that forecasts a result is not therefore a lever you can pull to change it.

The formal requirement makes the difficulty plain. For a surrogate to be trustworthy, the treatment’s effect on it must capture the treatment’s effect on the real outcome, a demanding condition that is rarely checked. And a sharper trap sits beneath it, the surrogate paradox: even when the surrogate and the outcome are strongly and positively correlated, a treatment that improves the surrogate can still worsen the outcome. A tight correlation between two measures does not license acting on one to move the other.

This is related to the familiar warning that a measure gamed into a target stops being a good measure, but it is not the same problem. A surrogate can fail when no one games anything, because the causal chain from proxy to outcome does not carry the whole of the intervention’s effect. Gaming is one way the link breaks; a treatment acting through unintended channels is another, and harder to see coming.

None of this is confined to medicine; it is the daily condition of program evaluation. The outcomes that justify our work are slow and costly, so we lean on proxies: placements, test scores, certificates, attendance, short-term output counts. Each is chosen because it correlates with the goal, and each can move without the goal following. A training program can lift placement rates by steering people into jobs that do not last; a school can raise scores by narrowing teaching to the test. The proxy improves, the report looks strong, and the mission is no better off.

The discipline is to keep the real outcome in view even when you cannot measure it directly. Name it explicitly, and treat the surrogate as a stated hypothesis about how you will reach it rather than proof that you have. Ask whether moving this proxy has actually moved the outcome before, for this kind of intervention. Look hard for the pathways by which your program could improve the proxy without helping, or while harming, the goal. And when the stakes are high, measure at least some true outcomes directly, on a subset or with longer follow-up, rather than trusting the proxy on faith.

A better number is easy to produce and easy to celebrate. Whether it means a better outcome is a separate question, and it is the one that finally matters.

So here is my question. For the proxies your program reports, can you point to evidence that moving them actually moves the outcome you care about, or are you trusting that the link will hold?

Posted in

Leave a comment