Every historian who reaches beyond a single case eventually confronts an unsettling question: on what grounds can we legitimately compare phenomena separated by centuries, languages, cosmologies, and institutional forms? The Roman familia and the modern nuclear family share a name in translation but almost nothing of their internal logic. Ming bureaucrats and Prussian civil servants both administered states, yet the very category of state imposes assumptions neither would have recognized.

This is the comparison problem in its sharpest form. Comparative history promises analytical leverage—the ability to distinguish contingent from structural, particular from general—but it purchases that leverage with categories that may distort the very objects it seeks to understand. The historian becomes a translator whose vocabulary was forged elsewhere.

Yet the alternative, radical incommensurability, threatens to dissolve history into a collection of isolated singularities, each intelligible only on its own terms and therefore intelligible to no one outside them. Between the Scylla of false universalism and the Charybdis of paralyzing relativism lies the actual practice of comparative historiography, which must justify itself philosophically even as it proceeds pragmatically. What follows examines the epistemological terrain of this dilemma and asks what disciplined comparison might look like.

Commensurability Questions

The strongest challenge to comparative history comes from what we might call the incommensurability thesis: that cultures constitute their own frameworks of meaning so thoroughly that concepts drawn from outside those frameworks cannot capture their internal logic. When we speak of religion in medieval Europe and religion in Tokugawa Japan, we deploy a category whose modern secular framing already prejudices the comparison.

This concern draws philosophical weight from thinkers as varied as Peter Winch, Clifford Geertz in his hermeneutic mode, and postcolonial critics who have shown how European analytic categories were universalized through conquest before being naturalized by scholarship. To compare, on this view, is often to smuggle in a tertium comparationis—a shared third term—that belongs to neither of the entities compared and secretly imposes its own conceptual architecture.

The problem intensifies when we compare across time. The past, as Bernard Bailyn observed, is a foreign country whose inhabitants cannot correct our misreadings. When we ask whether ancient slavery resembled modern slavery, we assume a stable referent for the term across two thousand years of shifting labor relations, legal codes, and moral vocabularies.

Yet the incommensurability thesis, pressed to its logical terminus, becomes self-undermining. To assert that two frameworks are incommensurable is already to have understood both sufficiently to compare their commensurability—a performative contradiction Donald Davidson identified decades ago. Complete conceptual isolation would make even the recognition of difference impossible.

The productive question is therefore not whether comparison distorts—all conceptualization involves selective abstraction—but which distortions we can render visible and account for. Incommensurability is better understood as a matter of degree, requiring careful calibration rather than wholesale refusal.

Takeaway

Radical incommensurability refutes itself: the very claim that two frameworks cannot be compared presupposes enough mutual intelligibility to establish the claim. The real work is calibrating the distortions comparison always introduces.

Comparative Gains

If the risks of comparison are real, so are its epistemic yields. Marc Bloch's foundational essay on comparative history argued that only through comparison can the historian distinguish what is genuinely specific to a case from what merely appears specific due to the narrowness of our vision. Single-case studies, however deep, cannot answer counterfactual questions about necessity and contingency.

Consider the debate over the great divergence—why sustained economic growth emerged in northwestern Europe rather than in the equally sophisticated economies of Song China or Mughal India. Only through systematic comparison did scholars like Kenneth Pomeranz and R. Bin Wong demonstrate that many previously assumed European advantages were shared or absent, redirecting explanation toward specific ecological and colonial contingencies.

Comparison also reveals patterns invisible from within any single tradition. The observation that agrarian empires across Eurasia developed strikingly similar solutions to problems of taxation, communication, and elite reproduction—despite minimal contact—suggests structural constraints that no purely internal history could detect. Similarities become data as much as differences do.

Perhaps most importantly, comparison denaturalizes the familiar. Studying Chinese examination culture illuminates the peculiarity of Western credential systems that seem to us self-evident. This estrangement effect, what Brecht called Verfremdung, is arguably comparison's most valuable philosophical yield: it returns our own historical formation to us as something requiring explanation rather than something taken for granted.

The gains, then, are not merely additive but transformative. Comparison changes what counts as a historical question, forcing us to explain what previously seemed to need no explanation.

Takeaway

Comparison's deepest value is not the discovery of similarities or differences but the estrangement effect: making our own tacit assumptions visible as historical peculiarities requiring explanation.

Methodological Standards

If comparison is neither impossible nor innocent, what disciplines it? A first principle is what Charles Tilly called the specification of units and their relations. Comparison requires clearly defined objects, and vague comparisons between civilizations or the West and the East almost invariably reproduce the ideological categories they claim to analyze.

A second principle involves the reflexive interrogation of one's tertium comparationis. When we compare, we must ask what conceptual apparatus makes the comparison possible and whether that apparatus derives asymmetrically from one of the cases being compared. This does not disqualify the comparison but requires that the asymmetry be acknowledged and its effects tracked.

Third, productive comparison distinguishes between what Reinhart Koselleck called synchronic and diachronic dimensions. Comparing societies at analogous developmental moments differs philosophically from comparing them at the same calendar date. Both are legitimate, but they answer different questions and rest on different assumptions about historical time itself.

Fourth, the comparative historian must specify what kind of claim comparison licenses. Comparisons can support claims about causal mechanisms, structural analogies, contingent similarities, or ideal-typical constructs in Weber's sense. Conflating these registers produces the false universalism that gives comparison its bad reputation.

Finally, comparison should remain revisable in light of what the compared cases reveal about the comparative framework itself. The best comparative work treats its categories as hypotheses tested against the recalcitrance of the material, not as fixed grids imposed upon it. This iterative discipline—between framework and evidence, between the general and the singular—is what separates comparison as inquiry from comparison as ideology.

Takeaway

Rigorous comparison requires treating one's own categories as hypotheses subject to revision by the cases they organize, rather than as neutral grids that stand outside the inquiry.

The comparison problem cannot be dissolved by methodological cleverness. It is a permanent condition of historical work whenever we reach beyond the single case, and even single-case work relies on implicit comparisons with the historian's own present.

What philosophical reflection offers is not escape from the problem but a more honest inhabitation of it. Comparison neither reveals timeless universals nor collapses into arbitrary imposition; it produces situated knowledge whose validity depends on the reflexive discipline of its practitioners.

The mature comparative historian works, then, in a mode Collingwood might recognize: reconstructing questions across contexts while acknowledging that the very act of reconstruction is itself historically situated. This is neither positivism's dream of neutral comparison nor postmodernism's refusal of it, but a chastened, self-aware practice that treats its own categories as historical objects while continuing to use them. The alternative—silence in the face of the world's variety—is not an option any serious historical discipline can afford.