Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

The Flaws of Quality Scores in Research Synthesis

Every systematic review rests on a step that rarely gets questioned: appraising the quality, or the risk of bias, of each study it includes. To do that, reviewers reach for an appraisal tool, a checklist or a scale. What almost no one asks is whether the tool itself is any good. An appraisal tool is a measurement instrument, and when you hold it to the same standards it imposes on the studies it judges, reliability and validity, it often falls short.

The oldest and most seductive form of appraisal is the summary quality score: run down a checklist, add up the points, and rank the studies by the total. That number feels objective and is anything but. It is a composite index, and it carries every problem a composite index carries, which items you include, how you weight them, how you combine them. The demonstration is now a classic: Jüni and colleagues scored the same trials with twenty-five different published quality scales and found that the choice of scale could change the conclusion of the meta-analysis, a trial ranked high by one scale ranked low by another. Sander Greenland had put it bluntly years earlier, that quality scores are useless and potentially misleading. The score launders a pile of judgment calls into one authoritative-looking figure, which is why the best modern tools refuse to produce one.

Even setting the score aside, appraisal is less reliable than its users like to think. Give the same study to two capable reviewers and they often disagree. The Newcastle-Ottawa Scale, the most widely used tool for observational studies, is the cautionary example: it has repeatedly shown poor agreement between reviewers, and it was never formally validated. That means a risk-of-bias rating is not a fixed fact about a study. It carries measurement error of its own, and a rating from a single reviewer is a shaky foundation to build a conclusion on.

There is a construct problem underneath the reliability problem. The idea of quality was always fuzzy, blending methodological rigor, completeness of reporting, and relevance, three different things. And an appraisal tool can only assess what the authors actually wrote down, so it partly measures how well a study was reported, not how well it was conducted. A careful study described carelessly gets marked down; a weak study written up smoothly can pass. The tool may be grading the prose as much as the science.

The deepest issue is that bias is not a single property of a study at all. It is specific to a mechanism and specific to an outcome. The clearest evidence comes from trials: inadequate concealment of who was assigned to which group, and lack of blinding, exaggerate treatment effects substantially for subjective outcomes but barely at all for objective ones like death. The same design feature is a serious flaw for one outcome and almost irrelevant for another in the very same study. This should sound familiar: just as a survey does not have one bias but a different bias for each estimate, a study does not have one quality. So the statement that a study is high quality is not well formed. The honest unit of appraisal is a specific bias, for a specific result.

The tool that grades the evidence deserves the same skepticism we bring to the evidence itself, and the single quality number is the least trustworthy thing the whole process produces. None of this means appraisal is hopeless. It means it has to be done differently, by domain, by outcome, and with its judgments made transparent rather than buried in a score. Its companion post takes up exactly that: if a single quality score fails, how should we appraise and use risk of bias instead?

So here is my question. When a review tells you a study is high or low quality, do you ask which bias, for which outcome, and how reliably that judgment was reached?

Posted in

Leave a comment