Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

Blog

Three people in old-fashioned pub discussing methodology

This site focuses on all things related to Research and Evaluation

My Latest Posts

• • •

  • You Are Not as Sure as You Think
    Try a quick test. Pick ten quantities you cannot know off the top of your head, the length of the Nile, the year a certain invention appeared, the population of… Read more: You Are Not as Sure as You Think
  • Beyond Averages: Designing for Diverse Outcomes
    In the late 1940s, United States Air Force jets were crashing with alarming frequency, and no mechanical fault could be found. Attention turned to the cockpit, which had been designed… Read more: Beyond Averages: Designing for Diverse Outcomes
  • Understanding the Surrogate Paradox in Health and Education
    We rarely measure the thing we actually care about. Durable employment, real learning, long-term health, and safety are slow to arrive, costly to observe, and hard to pin on any… Read more: Understanding the Surrogate Paradox in Health and Education
  • The Pitfalls of Retrospective Data Collection
    A large share of the data we collect asks people to remember. How many times did you see a doctor last year? How much did you drink last month? When… Read more: The Pitfalls of Retrospective Data Collection
  • Understanding the Limitations of Member Checking
    A common and well-meant step in qualitative research is to take your findings back to the people who gave you the data and ask, does this ring true? It feels… Read more: Understanding the Limitations of Member Checking
  • Identifying Bias in Linked Data Sets
    A great deal of modern research and evaluation runs on linked data. We connect a program’s enrollment file to earnings records, a survey to health claims, a benefits roster to… Read more: Identifying Bias in Linked Data Sets
  • Predictive vs Explanatory Modeling: Key Differences
    Two questions sound almost the same and are not. One is: what will happen? The other is: why does it happen, and what should we change? A model can be… Read more: Predictive vs Explanatory Modeling: Key Differences
  • Aligning Time to Avoid Immortal Time Bias
    Observational studies keep discovering that people who did a certain thing live longer. Patients who filled their prescriptions outlive those who did not. Heart transplant recipients outlive those on the… Read more: Aligning Time to Avoid Immortal Time Bias
  • Transforming Appraisal: From Scores to Structured Judgments
    Its companion post argued that a single quality score is unreliable and often invalid, and that bias is not one property of a study but something specific to a mechanism… Read more: Transforming Appraisal: From Scores to Structured Judgments
  • The Flaws of Quality Scores in Research Synthesis
    Every systematic review rests on a step that rarely gets questioned: appraising the quality, or the risk of bias, of each study it includes. To do that, reviewers reach for… Read more: The Flaws of Quality Scores in Research Synthesis
  • Beyond Response Rate: Measuring Survey Bias Correctly
    The first question people ask about a survey is almost always the response rate. Sponsors set targets for it, reviewers judge studies by it, and a low one is often… Read more: Beyond Response Rate: Measuring Survey Bias Correctly
  • The Impact of Anchoring on Estimates and Decisions
    Someone says a number out loud, a budget figure, a timeline, a rough guess, and from that instant your own estimate is quietly bent toward it. Not because the number… Read more: The Impact of Anchoring on Estimates and Decisions
  • The Ranking Is a Choice
    A single number that ranks things, states by vulnerability, hospitals by quality, programs by performance, carries enormous authority. It looks like a measurement, objective and settled. But a composite index is not a measurement the way a thermometer reading is. It is a construction, assembled through a chain of choices, and those choices, as much as the reality underneath, decide who ends up on top. Change one defensible step, the weighting, or the way the pieces are combined, and the order can reshuffle, sometimes dramatically. The danger is that these judgments disappear into a formula and the result arrives wearing the authority of arithmetic, right before someone makes a funding or accountability decision on it. Here is where the discretion hides, and how to read a ranking without being fooled by it.
  • Understanding Uncertainty in Data Visualization
    Picture a simple bar chart. Two bars, one taller than the other. Your eye settles the matter in an instant: this group is higher than that one. But the chart… Read more: Understanding Uncertainty in Data Visualization
  • Understanding the Table 2 Fallacy in Regression Analysis
    Open almost any study built on regression and you will find a table, often the second one, listing the outcome against a main variable and a row of controls, each with its own coefficient. The natural habit is to read down the column and treat every number as the effect of that variable. It is one of the most common mistakes in applied statistics, and it can quietly mislead a decision-maker about which factors matter. The model was built to answer exactly one causal question, and only one of its numbers is the answer; the rest can be biased in ways nobody checked, even when the main estimate is clean. Here is why, and how to read a regression table without falling for it.
  • Representativeness Is Not Always the Goal
    Ask most people how to choose a sample and they will reach for representativeness: draw at random, mirror the population, and bigger is better. For estimating a quantity in a population, that instinct is right. For a great deal of qualitative work, it is quietly wrong, and clinging to it leads people to dismiss good studies as too small or unrepresentative when neither is the real standard. In much qualitative research the goal is not to mirror a population but to understand something, and the best sample is the one that teaches you the most, which is often a handful of deliberately chosen cases, including the extreme and the atypical. Here is the logic of sampling for insight, and why the same tool that is dangerous in one setting is exactly right in another.
  • Two Ways to Be Uncertain
    You calculate a 95 percent confidence interval and describe it the natural way: there is a 95 percent chance the true value lies inside. Almost everyone reads it like this, and almost everyone is wrong, not because the math failed but because the sentence answers a question the tool was never built to answer. Behind a surprising share of misread statistics, the misused p-value, the overread confidence interval, is a quiet collision between two entirely different ways of thinking about uncertainty. Most of us unknowingly want the answer one of them gives while holding the tool of the other. Here is the difference between the two, why it explains so much confusion, and how to know which question your numbers are actually answering.
  • Not Every Program Holds Still – The Need for Developmental Evaluation
    Most of evaluation assumes a program that holds still: you specify the model, set the goals, let it run, then judge whether it hit them. That works when the program is stable and well understood. But a great deal of real work is not like that. New initiatives are still finding their shape, complex efforts sit in systems that keep shifting, and pilots are built to learn rather than to prove. Hold those against a fixed plan and you measure a moving target with a frozen ruler, and you may punish the very adaptation that makes them work. There is an approach designed for exactly this situation, and it asks a different question entirely, at a real cost worth naming. Here is what it is, and when it fits.
  • Survey Modes Matter: How Question Delivery Affects Responses
    You compare this year’s survey to last year’s, and the numbers have moved. Or you compare a phone sample to an online one, and they disagree. The natural reading is… Read more: Survey Modes Matter: How Question Delivery Affects Responses
  • The Risks of Misinterpreting Mediation Effects
    Knowing that a program worked is valuable. Knowing why it worked is more valuable still, because a mechanism you understand is one you can strengthen, cut, or carry to a… Read more: The Risks of Misinterpreting Mediation Effects
  • The Tradeoff Between Accuracy and Privacy in Data Analysis
    You pull a table from a federal statistical agency, or download a public microdata file, and you analyze it as though it were the unvarnished truth. It is not, and… Read more: The Tradeoff Between Accuracy and Privacy in Data Analysis
  • The Money Is Already Gone
    In my experience this is very common in our game. A program is failing and the evidence keeps piling up, but the argument for continuing never changes: we have already invested so much, we cannot stop now. It sounds like responsible stewardship. It is exactly backwards, and understanding why is one of the more useful corrections a decision-maker can make. Whatever you have already spent is gone regardless of what you do next, which is precisely what makes it the wrong thing to weigh. The same trap that keeps a failing program alive keeps a doomed pursuit on the board long after the odds have collapsed. Here is why we fall for it, what actually drives it, and how to decide the way the numbers say you should.
  • The Scale Ran Out of Room
    You run a solid program, measure the outcome before and after, and the scores barely budge. The obvious conclusion is that the program did not work. But there is a second explanation that has nothing to do with the program, and it is the kind of thing that quietly sinks evaluations: your measure may have had no room left to register improvement. When most people are already near the top of a scale, the instrument cannot show them getting better, and a real effect disappears into a flat result that looks like failure. The same trap can make two genuinely different groups look identical. Here is how it happens, how to spot it, and why a null result is sometimes a fact about your ruler rather than your program.
  • Your Sample Is Smaller Than It Looks
    You survey 1,500 people, and the number feels like strength: big sample, tight intervals, precise estimates. But suppose those 1,500 were not drawn independently. They were reached by picking 60 neighborhoods and interviewing 25 people in each. That one fact can cut the real precision of your survey by more than half, and if you analyze the data as though every row were independent, every standard error you report will be too small and every result will look more certain than it is. There is a single number that tells you how much information your clustered data actually holds, and for a lot of federal survey and program data, it is a sobering one. Here is what it is, and why it matters.
  • Causation Leaves Fingerprints
    Much of this series has been about establishing cause by comparison: a control group, a counterfactual, a case that did not get the treatment. But a great deal of real evaluation offers no such thing. There is one program, in one place, and nothing to compare it to. The usual instinct is to retreat into narrative and call it a story rather than a finding. There is a more disciplined option, and it comes from an unlikely place, the logic of a detective. It treats a causal claim as something that must have left evidence behind, and it has explicit rules for which evidence actually counts. Here is how a single case, examined rigorously, can still test a cause.
  • The Case Against P-Values: Rethinking Statistical Significance
    Over the course of this series I have returned, from several directions, to the same issue. This may be due to a bias that I picked up when I worked… Read more: The Case Against P-Values: Rethinking Statistical Significance
  • Why Small Studies Can Mislead Research Findings
    An underpowered study is usually described as one that might miss a real effect. That is true, and it is the least of the problem. The deeper danger is what… Read more: Why Small Studies Can Mislead Research Findings
  • Evaluating Data Fitness: Six Key Questions
    In my previous post from earlier today I argued that data is never good in the abstract, only fit or unfit for a particular use. That leaves the practical question.… Read more: Evaluating Data Fitness: Six Key Questions
  • Is Your Dataset Fit for Purpose?
    A dataset lands on your desk, large and clean, and someone asks the natural question: is it good? That question has no answer as posed. Data is not good or… Read more: Is Your Dataset Fit for Purpose?
  • The Shortcut in Every Survey
    A survey question asks a lot more than it appears to. To answer well, a respondent has to interpret what you meant, search memory for the relevant information, weigh it… Read more: The Shortcut in Every Survey

• • •