Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

Healthcare worker using AI-assisted screening showing false positive and demographic bias data

Few ideas in health feel more obviously correct than this one: catch the disease early, and you will do better. It is intuitive, it is often true, and it has driven decades of screening campaigns. So let me say clearly at the start that screening saves lives for several conditions, and nothing here argues for abandoning it. The argument is narrower and more uncomfortable: screening is not always beneficial, it can do net harm, and the reflex to detect everything earlier deserves more scrutiny than it usually gets.

Start with why early detection helps only sometimes. Finding a disease sooner changes the outcome only if acting sooner changes its course. For some diseases it does. For others, the timing of detection does not change where the story ends, and moving the diagnosis earlier only means the person lives longer as a patient, not longer overall.

That hides a measurement trap that makes screening look better than it is. Three biases inflate its apparent benefit:

1. Lead-time bias. Move the diagnosis date earlier and you lengthen survival measured from diagnosis, even if the date of death does not move at all. People appear to survive longer when all that changed was the starting line.

2. Length-time bias. Screening preferentially catches slow-growing cases, because they sit in the detectable but silent stage far longer. The fast, aggressive cases tend to surface as symptoms between screening rounds. So screen-detected disease is a gentler, more survivable sample of the disease to begin with.

3. Overdiagnosis. The extreme of length bias. Screening detects disease that never would have caused symptoms or death in the person’s lifetime. These patients cannot be helped, because there was nothing to prevent. They can only be harmed by the treatment that follows.

This is why counting cancers caught early, or comparing survival rates, is the wrong yardstick. Both are inflated by the biases above. The honest measure is whether screening lowers mortality, ideally death from all causes, in a fair comparison. Optimizing the easy proxy, cases found, instead of the real goal, deaths prevented, is the same trap this series flagged with Goodhart’s law.

The false-positive problem compounds this, and it is a base-rate problem in disguise. When a disease is rare, even an accurate test produces many more false alarms than true findings, so a positive result can mean far less than it seems. Across repeated rounds of screening, the chance of at least one false positive climbs steadily, and each one can bring anxiety, follow-up procedures, and biopsies that carry their own risks. False negatives, in turn, deliver false reassurance.

The cautionary cases are real and well documented. When South Korea added thyroid ultrasound to routine screening, thyroid cancer diagnoses rose roughly fifteenfold over about two decades while deaths from thyroid cancer stayed flat, a textbook overdiagnosis epidemic that produced a wave of thyroid surgeries, each with real complications. Infant screening for neuroblastoma, tested in controlled studies in Germany and Canada, raised detection substantially but did not reduce advanced disease or mortality, and Japan ended its national program. The long debates over PSA testing and over mammography turn on these same trade-offs, which is why expert bodies have moved toward individualized decisions and still disagree about exactly when and how often to screen.

None of this means early detection is a myth. It means screening is a medical intervention like any other, with benefits and harms to be weighed for each disease and test, not assumed. The classic checklist, from Wilson and Jungner at the World Health Organization in 1968, still holds: the condition should matter, there should be a recognizable early stage, the test should be acceptable, treatment should work better when started earlier, and the benefits should outweigh the harms and costs. Cervical and colorectal screening clear that bar convincingly. Others clear it only for certain ages or risk groups, and a few do not clear it at all.

So yes, we should temper the demand. Early detection is a means, not an end. The end is less suffering and fewer deaths, and screening earns its keep only when it delivers them. The mature position is not screen everything or screen nothing. It is to ask, honestly and disease by disease, whether finding this earlier actually helps, and to tell people the whole story, including the harms, not only the reassuring half.

So here is my question. When you see a screening program promoted with survival rates and cases caught early, do you ask the harder question of whether it actually lowers mortality, and at what cost?

Posted in

Leave a comment