Stephen T’s Blog Spot

A blog aimed at issues only data scientists, data analysts, statisticians, evaluators, and researchers care about.

Transforming Appraisal: From Scores to Structured Judgments

Its companion post argued that a single quality score is unreliable and often invalid, and that bias is not one property of a study but something specific to a mechanism and an outcome. That is the diagnosis. The repair is not a better scale. It is a different way of working, one that treats appraisal as transparent, structured judgment rather than a number.

Start by refusing to sum. The modern tools are built this way on purpose: the Cochrane risk-of-bias tool for randomized trials, and ROBINS-I for non-randomized studies, ask for a separate judgment in each bias domain, how participants were assigned, whether groups stayed comparable, how missing data were handled, how outcomes were measured, whether results were selectively reported, and keep those judgments apart instead of collapsing them into one figure. A study becomes a profile across domains, not a rank, which removes the arbitrary weighting that made summary scores indefensible.

Not every domain carries equal evidence, so weight your attention accordingly. A few items have been shown empirically to predict distorted effects, concealment of allocation and blinding chief among them, and the meta-epidemiological studies tell you not just that they matter but when. Take those seriously. Be honest that other items on the checklist rest on expert consensus rather than demonstrated bias, and do not give a consensus item the same evidentiary weight as one with a track record.

Because the same flaw bites differently depending on what is measured, appraise risk of bias for each outcome, not once for the whole study. A trial can be at low risk for an objective outcome like mortality and higher risk for a subjective one rated by unblinded assessors. Rating the study once, and stamping that rating on every result it reports, throws away the very distinction that matters most.

Since two reviewers will disagree, manage it the way you would manage any unreliable measurement. Use two independent appraisers, write down decision rules before you start, pilot the tool on a few studies to calibrate, and record the reasoning behind each judgment so a reader can check it. The goal is not to eliminate judgment, which is impossible, but to make it transparent and reproducible.

Here is the move that matters most, and it is one this series has made before in another setting. Do not use risk of bias as a gate that admits or excludes studies, and do not fold it into a weight. Use it as a sensitivity analysis: pool the evidence, then ask what happens when you restrict to the studies at lowest risk of bias. If the effect holds, the conclusion is robust. If it shrinks toward nothing once the most biased studies are removed, the bias was doing the work, and that is itself the finding. Risk of bias earns its keep as a question you ask of the result, not a filter you apply to the inputs.

At the level of the whole body of evidence, this is what GRADE formalizes, rating certainty for each outcome by combining risk of bias with consistency, directness, precision, and the threat of publication bias. Used well, it communicates how much to trust a result. Used carelessly, its tidy structure can lend an unearned air of objectivity to a chain of judgment calls, so the discipline is to show the reasoning, not just the rating.

The through-line of both posts is the same. The aim of appraisal is not a number that ranks studies but a transparent, domain-by-domain, outcome-specific judgment you can defend and stress-test. Grade the bias, for the outcome, out loud, and let the conclusion prove it can survive the least biased evidence.

So here is my question. In your reviews, does risk of bias end up as a score that ranks studies, or as a question you put to the result to see whether it holds?

Posted in

Leave a comment