In 1915, when Einstein completed his general theory of relativity, he already knew that Mercury's perihelion precessed by an anomalous 43 arcseconds per century. This wasn't a prediction discovered through his equations—it was a puzzle astronomers had wrestled with since 1859. Yet when Einstein's field equations reproduced that exact figure, physicists reacted as though something profound had been confirmed.
Here lies a philosophical puzzle that has haunted the theory of scientific inference for decades. If confirmation is fundamentally about a theory raising the probability of its evidence, how can evidence you already possess—evidence with probability effectively equal to one—raise the probability of anything at all? The Bayesian machinery seems to grind to a halt precisely where scientific practice hums along most confidently.
This is Clark Glymour's problem of old evidence, and it reveals a peculiar tension between how philosophers model scientific reasoning and how scientists actually deploy it. When Darwin marshalled decades of accumulated natural history observations to argue for descent with modification, was he doing something epistemically inferior to a scientist who predicts novel phenomena? Kuhn suggested that revolutions succeed partly by explaining stubborn anomalies already known. If that explanation confers no confirmatory weight, we have a problem—not with science, but with our theory of it.
The Temporal Paradox
The paradox begins with an innocent-looking axiom of confirmation theory. Evidence E confirms hypothesis H when the probability of H given E exceeds the prior probability of H. This Bayesian formulation, elegant in its simplicity, requires that E was previously uncertain. Otherwise, conditioning on E changes nothing, and confirmation vanishes.
But scientific reasoning routinely violates this cleanness. By the time general relativity emerged, Mercury's anomalous perihelion had probability one for every working astronomer. It was as certain as any empirical fact could be. Formally, then, it could not confirm Einstein's theory. Yet it manifestly did—the calculation was celebrated as a triumph, cited in Nobel deliberations, and remains textbook evidence for the theory a century later.
The tension isn't merely technical. It touches something deep about the temporal structure of inquiry. Theories are formulated in a historical context saturated with prior observations. Newton knew about Kepler's laws before writing the Principia. Mendeleev knew the properties of dozens of elements before constructing his periodic table. If old evidence cannot confirm, then a vast portion of what scientists take themselves to be doing is philosophically illegitimate.
Several strategies have been proposed to dissolve the paradox. Some suggest we should imagine a counterfactual epistemic state in which we didn't yet know E—reasoning about what our credences would have been. Others argue that what genuinely gets confirmed is the logical connection between theory and evidence, discovered only when the theory is formulated. On this view, the news isn't the perihelion's precession—it's that general relativity entails it.
Neither move is entirely satisfying, and the debate continues. What matters for practicing scientists is recognising that the puzzle is real, and that the intuitive weight we assign to a theory's ability to account for known anomalies deserves careful philosophical scrutiny rather than uncritical acceptance.
TakeawayThe most powerful confirmations in scientific history often involve evidence that was already sitting on the table. If your epistemology can't explain why, it's your epistemology that needs revision, not the history.
Prediction Versus Accommodation
The old evidence problem opens onto a larger controversy: does successful prediction carry more evidential weight than after-the-fact accommodation? Popper insisted it must. Risky predictions that could have failed but didn't, he argued, place a theory in genuine jeopardy. Accommodation, by contrast, merely fits the theory to known facts—a curve drawn through pre-existing points.
This predictivist intuition runs deep. When Halley's comet returned in 1758 as Newtonian mechanics predicted, or when Eddington measured stellar deflection in 1919, something seemed epistemically special about the temporal ordering. The theory stuck its neck out. Nature could have refused to cooperate. That it didn't feels like stronger vindication than merely recovering what we already knew.
Yet accommodationists press back with formidable arguments. From a Bayesian standpoint, the confirmatory power of evidence depends on the likelihood ratio—how probable the evidence is given the theory versus its negation—not on when the evidence was collected. The perihelion of Mercury has the same probability under general relativity whether that probability is computed in 1859 or 1915. Time should be epistemically inert.
The resolution, many now argue, lies in unpacking what predictivism was really tracking. When a theory predicts novel evidence, it typically wasn't designed with that evidence in mind, so it couldn't have been fudged to fit. Accommodation, by contrast, permits ad hoc tuning—parameters adjusted precisely to accommodate known facts, sacrificing genuine explanatory content for surface agreement. The relevant epistemic distinction, then, isn't temporal but structural: was the theory constructed to fit the evidence, or did the fit emerge as a genuine consequence?
This reframing helps explain why Einstein's Mercury calculation felt so significant. General relativity wasn't crafted to explain the perihelion. The 43 arcseconds fell out of equations motivated by entirely different considerations—equivalence principles, general covariance, the geometry of spacetime. The old evidence became new confirmation because it emerged unforced from independent theoretical commitments.
TakeawayWhat matters epistemically isn't whether evidence came before or after the theory, but whether the theory could have been quietly tuned to fit it. Genuine confirmation requires that the fit wasn't engineered.
How Scientists Actually Use Old Evidence
Set aside the philosophical puzzles for a moment and observe how researchers actually reason. In practice, scientists constantly deploy previously known facts to evaluate new theories, and they do so with sophisticated implicit standards that partly track the philosophical distinctions above.
Consider theoretical unification. When Maxwell showed that his electromagnetic equations implied light must travel at a specific velocity—matching the empirically measured speed of light—no new measurement was made. Yet the coincidence transformed physics. The old evidence about light's speed became newly meaningful because it emerged from theoretical structure built for entirely different purposes: unifying electricity and magnetism. Scientists intuitively recognise that when disparate domains converge on the same numerical value, coincidence becomes vanishingly implausible.
In evolutionary biology, Darwin's argument in The Origin is essentially a masterwork of old-evidence reasoning. He didn't predict novel phenomena so much as demonstrate that an enormous, disparate body of facts—biogeography, embryology, comparative anatomy, the fossil record—coheres under a single explanatory framework. Modern biologists still cite this integrative power as central to evolution's evidential standing.
The practical criterion scientists seem to employ resembles what philosophers call use-novelty: what matters isn't whether evidence was known before the theory, but whether it was used in the theory's construction. If a theorist built her framework without consulting a particular dataset, and the framework nonetheless recovers that dataset accurately, this counts as genuine confirmation—regardless of temporal ordering.
This tacit standard explains why fields have developed practices like preregistration and blinded analysis. These procedures aren't merely bureaucratic hygiene. They create conditions under which accommodation cannot masquerade as prediction, protecting the epistemic status of evidence that might otherwise be quietly incorporated into the very theories it purports to test.
TakeawayScientific practice has quietly resolved what philosophy still debates: evidence confirms theories when it emerges unforced from independent theoretical commitments, not when it happens to arrive in the right chronological order.
The problem of old evidence begins as a technical wrinkle in Bayesian confirmation theory but ends by illuminating something essential about scientific reasoning. Confirmation isn't merely about temporal surprise—it's about whether the fit between theory and world was earned or engineered.
This reframing carries practical weight. It suggests that theories which unify previously disparate observations deserve substantial credit, even when no novel prediction is involved. It also warns against the seductive appearance of accommodation, where clever theorists can quietly tune their frameworks to any pattern they choose to explain.
Perhaps the deepest lesson is methodological humility. Our formal models of scientific inference remain incomplete, unable fully to capture practices that working scientists execute intuitively. The old evidence problem reminds us that philosophy of science, at its best, doesn't legislate to inquiry—it strives to understand what inquiry has already discovered how to do.