Category: Uncategorized
-
You pull a table from a federal statistical agency, or download a public microdata file, and you analyze it as though it were the unvarnished truth. It is not, and the gap is deliberate. Before that data reached you, someone changed it to protect the people in it. They may have suppressed small cells, swapped…
-
In my experience this is very common in our game. A program is failing and the evidence keeps piling up, but the argument for continuing never changes: we have already invested so much, we cannot stop now. It sounds like responsible stewardship. It is exactly backwards, and understanding why is one of the more useful…
-
You run a solid program, measure the outcome before and after, and the scores barely budge. The obvious conclusion is that the program did not work. But there is a second explanation that has nothing to do with the program, and it is the kind of thing that quietly sinks evaluations: your measure may have…
-
You survey 1,500 people, and the number feels like strength: big sample, tight intervals, precise estimates. But suppose those 1,500 were not drawn independently. They were reached by picking 60 neighborhoods and interviewing 25 people in each. That one fact can cut the real precision of your survey by more than half, and if you…
-
Much of this series has been about establishing cause by comparison: a control group, a counterfactual, a case that did not get the treatment. But a great deal of real evaluation offers no such thing. There is one program, in one place, and nothing to compare it to. The usual instinct is to retreat into…
-

Over the course of this series I have returned, from several directions, to the same issue. This may be due to a bias that I picked up when I worked at ECRI writing systematic reviews for AHRQ’s Evidence Based Practice Center program. I read an article, authored by Sander Greenland and colleagues, on the worth…
-

An underpowered study is usually described as one that might miss a real effect. That is true, and it is the least of the problem. The deeper danger is what happens when a small, noisy study does find something. A statistically significant result from an underpowered study is probably a large overestimate of the true…
-

In my previous post from earlier today I argued that data is never good in the abstract, only fit or unfit for a particular use. That leaves the practical question. Given a dataset and a purpose, how do you actually decide whether the first can serve the second? The answer is not a feeling about…
-

A dataset lands on your desk, large and clean, and someone asks the natural question: is it good? That question has no answer as posed. Data is not good or bad in the abstract. It is fit, or unfit, for a particular use. The same dataset can be excellent for one question and worthless, even…
-

A survey question asks a lot more than it appears to. To answer well, a respondent has to interpret what you meant, search memory for the relevant information, weigh it into a judgment, and then map that judgment onto the response options you offered. That is real cognitive work, and we assume that respondents are…
-

In a previous post I argued that effectiveness is only half the question, and that a program can work and still be a poor use of money. That raises the obvious follow-up: how do you actually build a cost-effectiveness case that will survive scrutiny? The good news is that most of it is not exotic…
-

An evaluation comes back positive. The program works, the effect is real, and the instinct is to call it a success and recommend expanding it. But whether it works is only half of the question a decision-maker actually faces. The other half is whether it is worth it, and a program can clear the first…
-

You build a model and it fits beautifully, and that is often a warning, not a triumph. A model flexible enough to fit every wrinkle in your data is fitting noise as well as signal, and the noise will not repeat, so it predicts new data worse. This is overfitting. Because in-sample fit is inflated,…
-

On the Fourth of July, lets take a detour and look at the Declaration of Independence as an argument about evidence. Before listing any grievance, it sets a standard: a decent respect to the opinions of mankind requires declaring your reasons. Then it promises to let facts be submitted to a candid world, and delivers…
-

Ask an experienced team how long a project will take, and they will study the specifics. They will break the work into phases, estimate each one, add the pieces, and perhaps pad the total for safety. It is a careful, disciplined process, and it produces an answer that is almost always too optimistic. The very…
-

Every causal claim from observational data rests on an assumption that cannot be checked. When you estimate the effect of a program, a treatment, or a policy from data you did not randomize, you are assuming that you have measured and adjusted for every important confounder. There is no test for this. The variable that…
-

You field a survey, score it, and compare two groups. Men score higher than women on the scale, or one site outperforms another, or the average climbs after the program. The natural next move is to interpret the difference. But there is a prior question that almost no one asks, and it can dissolve the…
-

In the previous post I argued that screening can do harm, and that the usual evidence offered for it, more cases caught and higher survival rates, is exactly the evidence that misleads. That naturally raises the next question. If survival and cases found are the wrong measures, what are the right ones? How do you…
-

Early disease detection is generally seen as beneficial, as it can save lives through screening. However, screening does not always result in better health outcomes and may even cause harm. The effectiveness of early detection depends on the disease and whether early intervention alters its progression. Biases, like lead-time and length-time biases, can inflate perceived…
-

“Does the program work?” sounds like the most basic question an evaluator can ask. It is also close to unanswerable as written, because the same program routinely succeeds in one place and fails in another, for reasons that a simple yes or no can never hold. The honest answer is almost always: it depends. Realist…