A child reads at the 60th percentile. An applicant scores 110 on an aptitude scale. A patient lands in the average range on a cognitive screen. These numbers feel like objective readings, the way a thermometer reports a temperature. They are not. Each one is a statement about where the person stands relative to a particular group of other people, the norming sample, measured at a particular time. Change the reference group and the same performance earns a different score.
Most standardized scores are norm-referenced, and it is worth being clear about what that means. The test was given to a reference sample, and your score is your rank within that sample’s distribution. A percentile reports what fraction of that group you outscored; a standard score places you in standard-deviation units around that group’s average. The number does not describe what you can do in absolute terms. It describes where you sit in a crowd, and the crowd is doing the defining.
That leads to the first problem: the reference sample may not be your population. A norm is only as good as the sample it was built on. If the norming group was drawn from people unlike those you are testing, a different country, era, income level, or amount of schooling, the percentiles are calibrated to the wrong crowd, and your group’s scores shift systematically. A cutoff normed on one population can flag far too many people in another, or far too few. This is coverage error, from an earlier post, living inside a single number.
The second problem is that norms age. Reference distributions drift as the population changes. The famous case is the Flynn effect: across the twentieth century, raw performance on cognitive tests rose by roughly two to three points per decade, so a test normed decades ago compares today’s takers against a weaker, older reference group and inflates their scores. Recent data show the drift can slow or even reverse in places, which only sharpens the lesson: a norm is a snapshot of a moving population. A score read against expired norms is compared to people who no longer exist, which is why serious tests are renormed and why scores from different editions are not interchangeable.
The danger concentrates at decision thresholds. When a fixed cutoff governs eligibility, the same raw performance can land above the line on old norms and below it on current ones. In the diagnosis of intellectual disability, where the threshold carries clinical and even legal weight, the choice of norm can move a person across the boundary. When a number decides something, whose yardstick you used stops being a technicality.
It helps to see that there is another way to score entirely. A criterion-referenced measure asks whether you can do a defined thing, meet this standard, clear this bar, regardless of how anyone else performed. Norm-referenced scoring answers how you compare with others; criterion-referenced scoring answers whether you can do the task itself. Both are legitimate; however, they answer different questions, and reporting one as though it were the other is a common error. A percentile cannot tell you whether a student has mastered the material, and a mastery standard cannot tell you how rare that mastery is.
None of this is far from our daily work. Evaluation and selection lean on standardized scores: reading levels, aptitude and certification results, screening cutoffs, risk indices. When we compare a program’s participants to a national norm, or draw a cutoff from a normed instrument, we inherit whoever was in that norming sample and whenever it was collected. A gain that looks real may be partly the drift of stale norms; a disparity may be an artifact of a reference group that never resembled the people served. Before trusting a standardized score, ask whose distribution it is scored against, and when.
So here is my question. For the standardized scores you rely on, do you know who was in the norming sample and how long ago they were measured, or are you reading a relative rank as if it were an absolute fact?
Leave a comment