Consider a puzzle that has vexed economic historians for decades: literacy rates in early modern Europe, estimated from signatures on marriage registers, suggest a steady climb from roughly 30 percent in 1600 to 65 percent by 1800 in parts of England. Yet when researchers cross-reference these figures with probate inventories and tax records, systematic distortions emerge. The signatories were not a random cross-section of the population—they were the marrying, the surviving, the recorded.
This is the central methodological hazard of quantitative history: our data does not arrive by random sampling. It arrives through filters—of survival, selection, and observation—each of which introduces bias that can invert our conclusions. A dataset of medieval merchant accounts tells us about merchants whose accounts survived, kept by clerks whose employers prospered, in cities whose archives escaped fire and war.
The problem is not that historical data is bad. It is that historical data is selected, often in ways that correlate directly with the phenomena we wish to measure. Growth, wealth, mobility, and mortality all interact with the mechanisms that determine what evidence reaches us. Ignoring this generates what statisticians call endogenous sampling—and produces confidently wrong answers. What follows examines three dimensions of this problem and the econometric tools developed to address them.
Survival Bias: The Silent Editor of the Historical Record
Documents do not survive at random. They survive because someone valued them, stored them properly, or because chance spared them from fire, flood, and neglect. This means the surviving corpus is systematically correlated with wealth, institutional continuity, and geographic stability—precisely the variables historians often wish to measure.
Consider Gregory Clark's work on English real wages. Early estimates drawing on institutional wage series—Oxford colleges, cathedral chapters, royal building projects—suggested relatively stable living standards from 1300 to 1800. But institutional employers preserved records because they were institutions. Their wage rates were sticky, formalized, and unrepresentative of the volatile day-labor market where most workers actually earned their bread.
The magnitude of this distortion can be staggering. When Robert Allen reconstructed European wages using broader sources including private accounts, ecclesiastical fabric rolls, and municipal records, the picture shifted substantially. Estimated real wages in 17th-century London rose by roughly 20 percent relative to the institutional-only series. The direction of some long-run trends reversed.
Survival probability typically correlates with three factors: institutional backing (churches and universities outlast individuals), geographic advantage (dry climates and stable regions preserve better), and social status (elite records were considered worth preserving). A naive analysis treats survival as random. A rigorous one treats it as a selection equation.
The practical implication is that any historical dataset should be accompanied by an explicit model of its generating process. What produced this record? Why did it survive? What comparable records did not survive, and how did they systematically differ? These are not rhetorical questions—they are prerequisites for valid inference.
TakeawayThe historical record is not a random sample of the past; it is a curated exhibit assembled by the biases of survival. What we can measure is always shaped by what chose to endure.
Migration Selection: When Movement Rewrites the Sample
Historical demographers face a persistent challenge: those who move differ systematically from those who stay. This truism has profound consequences for how we interpret data on wages, health, fertility, and social mobility. Any cross-sectional snapshot of a population reflects not the population's underlying characteristics, but a selected subset shaped by migration decisions.
The classic case is the analysis of 19th-century transatlantic migration. Early studies comparing wages of Irish immigrants in Boston with those who remained in Ireland concluded that migration produced modest earnings gains. But Joseph Ferrie's linked-census work revealed that emigrants were positively selected on unobservable traits—ambition, health, literacy—that would have raised their earnings anywhere. Controlling for selection cut estimated migration premiums by roughly a third.
The direction of selection is not always positive. During famines, the migrants may be those with just enough resources to leave; the destitute stay and die. During prosperity, migrants may be the marginally employed rather than the successful. The selection mechanism varies with context, and assuming a fixed direction produces systematic error.
This becomes acute when studying long-run outcomes. A community's measured characteristics in 1850 reflect who arrived, who left, and who died since 1800. If out-migrants were disproportionately young and healthy, the community appears to have aged and grown sicker even if no individual changed. Aggregate trends can be entirely artifactual.
Modern approaches use linked longitudinal records—following named individuals across censuses, ship manifests, and vital records. This transforms a repeated cross-section into a panel, allowing researchers to distinguish compositional change from behavioral change. The methodological cost is enormous; the inferential gain is transformative.
TakeawayPopulations are not static containers of data—they are shaped continuously by who arrives, departs, and dies. Any snapshot conflates behavior with sorting.
Correction Methods: Recovering Signal from Selected Samples
The recognition of selection bias would be intellectually paralyzing if not accompanied by tools to address it. Fortunately, econometrics offers a toolkit—developed largely since Heckman's foundational 1979 paper—that allows partial correction of biased samples, provided the researcher can model the selection process itself.
The Heckman two-stage estimator remains the workhorse. In stage one, a probit model estimates the probability of selection—of appearing in the sample—as a function of observable variables. In stage two, the inverse Mills ratio derived from this probability enters the outcome regression as an additional regressor, absorbing the bias from non-random selection. Applied to historical data on female labor force participation, apprenticeship completion, or literacy, this technique can substantially revise naive estimates.
For survival-biased samples, the analogue is often a truncation model or a Tobit specification, where the researcher explicitly models the threshold determining whether a record enters the dataset. Manuscript census evaluations by Steckel and others use these methods to correct for underenumeration of infants, young men, and marginal populations.
Where formal models are unavailable, bounding approaches offer a more agnostic alternative. Manski's partial identification framework asks: given the range of plausible selection mechanisms, what is the range of possible true parameter values? This yields intervals rather than point estimates but avoids the false precision of models resting on untestable assumptions.
The essential discipline is transparency. Every correction method rests on assumptions about the selection process that cannot be fully verified from the data itself. Robust research reports naive estimates, corrected estimates, and sensitivity analyses showing how conclusions change under alternative selection models. Precision without honesty about assumptions is worse than useful ignorance.
TakeawaySelection bias cannot be eliminated, but it can be modeled, bounded, and disclosed. The mark of rigorous quantitative history is not the absence of bias but its explicit accounting.
The quantitative turn in history has delivered remarkable insights, but its credibility rests on how honestly it confronts the non-random nature of historical evidence. Every dataset is the endpoint of a long selection process—of survival, of migration, of documentation—and treating it otherwise risks confidently reproducing the biases embedded in the archive.
The most productive research programs going forward will combine three elements: explicit modeling of selection mechanisms, use of linked longitudinal records where possible, and transparent sensitivity analysis showing how conclusions depend on identifying assumptions. This is more demanding than treating data as raw truth, but the additional rigor is what distinguishes robust findings from statistical folklore.
Further work is needed on bounding methods for historical inference, on machine-learning approaches to record linkage, and on the systematic auditing of foundational datasets. Selection bias is not a solved problem—but the tools to think clearly about it now exist. The question is whether the field will consistently use them.