A large share of the data we collect asks people to remember. How many times did you see a doctor last year? How much did you drink last month? When did the symptoms start? Since the program ended, have you found work? We treat the answers as records of what happened. They are not records. They are reconstructions, and memory reconstructs with a systematic tilt.
The mind does not store events like a tape and play them back on request. It rebuilds them from fragments, cues, and present beliefs, and that rebuilding introduces error. The error is not random noise that averages out; it follows predictable patterns that push estimates in particular directions. Knowing the patterns separates a usable retrospective measure from a misleading one.
Several patterns do most of the damage. Recall decay is the simplest: the further back you ask, the more is forgotten, so long recall periods undercount events, especially minor ones. Telescoping is sneakier: people misdate events, usually pulling them forward so they feel more recent, which inflates the count inside a bounded window like the past twelve months, because older events get dragged into it. Rounding and heaping pile responses onto salient numbers, zero, five, ten, about twice a week, distorting the distribution and its tails. And salience skews what survives at all, since vivid or emotional events are remembered while routine ones vanish, so rare dramatic behaviors are overcounted and frequent mundane ones undercounted.
Then there is the pattern that turns error into bias. Current state colors memory: people who feel unwell now recall more past symptoms, and people who believe something helped recall their earlier situation as worse than it was. This is harmless as long as it operates equally across the groups you compare; it becomes dangerous the moment it does not. In a case-control study, people who have the disease search their memory harder for causes than healthy controls do, so the cases recall more exposure and an association appears that may not be real. This is differential recall, and it does not merely add noise; it manufactures or erases the very difference you are trying to measure.
Evaluation has its own popular version. Asking people at the end of a program to rate where they were before it, thinking back, how confident were you at the start, is cheap and sidesteps some real problems with baseline surveys. However, it hands the respondent’s current beliefs about the program a direct channel into the baseline measure. If they think the program worked, they will tend to remember a lower starting point, and the design will produce an effect whether or not one occurred. A retrospective baseline is the easiest way to measure a program’s reputation and call it impact.
The repair follows the diagnosis. Prefer contemporaneous measurement whenever you can afford it: capture data at baseline, use diaries or real-time prompts, or draw on records instead of memory. When you must ask retrospectively, shorten and bound the recall window, and anchor dating with calendars or landmark events to blunt telescoping. Offer specific options to recognize rather than forcing people to generate answers from nothing, and validate self-report against records where any exist. Above all, match recall conditions across your groups, so that any recall error is similar on both sides rather than concentrated in the group with the most reason to reconstruct the past.
None of this is academic for program evaluation, which runs on retrospective self-report: exit surveys, follow-up interviews, questions that begin with since the program. The convenience is genuine; however, the recall structure quietly favors finding an effect, because the people most invested in a program are often the ones whose memory of before has been most reshaped by after. Treating memory as a reconstruction, whose errors are patterned and can differ by group, is the difference between measuring what a program did and measuring what people now believe about it.
So here is my question. When your data depend on what people remember, do you account for the fact that memory is reconstructed and skewed, and that the skew may be larger in exactly the group you expect to benefit?














