When we speak of a tree, we compress an astronomical amount of physical detail—cellulose fibers, xylem channels, photosynthetic gradients—into a single, portable concept. This compression is not arbitrary. Something about the structure of reality invites us to draw a boundary around this particular pattern of matter and treat it as a unit. But is that boundary discovered or invented?

The question becomes urgent when we consider the artificial minds now learning representations of the world at a scale that dwarfs any single human. If sufficiently capable systems inevitably carve nature at the same joints we do, then a hidden bridge exists between us and them—a shared conceptual substrate that could ground genuine understanding. If they do not, we may find ourselves conversing with intelligences whose inner ontologies bear no resemblance to our own, even when their outputs superficially mirror our language.

This is the terrain mapped by the natural abstractions hypothesis, a claim about the structure of information in complex systems that carries profound implications for alignment, interpretability, and the future of human-machine coexistence. It asks whether the concepts we take for granted are contingent artifacts of primate cognition or discoverable features of the world itself—and whether the minds we are building are our epistemic cousins or something altogether stranger.

The Hypothesis: Reality's Preferred Coordinates

The natural abstractions hypothesis, articulated most rigorously by John Wentworth, proposes that the universe contains privileged summaries. Information about a system's fine-grained state, when propagated across sufficient distance or through enough noisy intermediaries, tends to be destroyed—except for certain low-dimensional features that persist. These surviving features are what any far-away observer, whether human or otherwise, is forced to use in order to predict the system's behavior.

Consider a cloud of gas. Its precise molecular configuration is inaccessible to a distant observer, but its temperature, pressure, and volume propagate outward robustly. These are not arbitrary human categories; they are the information bottlenecks that survive coarse-graining. The claim generalizes: chairs, cats, and causes may be natural in the same sense—stable summaries that any sufficiently powerful predictor must converge upon.

If true, this reframes abstraction as discovery rather than invention. Concepts would be less like arbitrary linguistic conventions and more like periodic table elements—features of reality waiting to be found. A superintelligent AI, on this view, would not construct a wholly alien conceptual scheme. It would encounter the same joints in nature that human cognition has painstakingly located over evolutionary and cultural time.

The mathematical scaffolding involves information theory and the redundancy of causal structure. Wentworth's formalism suggests that in systems with sufficient hierarchy and locality, the abstractions useful for prediction are objectively determined by the system itself, not by the observer's contingent history. This is a strong claim, and its plausibility depends on assumptions about the world's structure that remain empirically contested.

Yet the hypothesis is not merely philosophical speculation. It generates testable predictions: representations learned by diverse architectures on diverse tasks should exhibit systematic overlap in what they encode. Failure of this convergence would falsify the strong version; confirmation would suggest we share more with our creations than intuition might grant.

Takeaway

Abstractions may be less like human inventions and more like fossils—patterns pressed into the structure of information itself, waiting for any sufficiently curious mind to unearth.

Evidence from the Interpretability Frontier

Empirical work on neural network interpretability offers tantalizing but ambiguous evidence. Anthropic's research on sparse autoencoders has revealed that large language models contain identifiable features corresponding to concepts remarkably close to human-legible categories: the Golden Gate Bridge, code vulnerabilities, sycophancy, deception. When probed carefully, the model's internal geometry seems to carve territory in ways we can recognize.

More striking is the phenomenon of cross-model convergence. Studies comparing representations across independently trained networks—different architectures, different datasets, different initialization seeds—find surprising alignment in the features learned at intermediate layers. The Platonic Representation Hypothesis, recently proposed by researchers at MIT, argues that as models scale, their representations converge toward a shared statistical structure of reality.

Vision and language models trained separately develop internal spaces that can be aligned with linear transformations, suggesting they have found overlapping compressions of the same underlying world. This is not conclusive proof of natural abstractions, but it is consistent with the hypothesis in a way that would be surprising if concept-formation were arbitrary.

Counter-evidence exists as well. Adversarial examples reveal that models rely on features imperceptible and unintuitive to humans—textures and statistical regularities that carry predictive power but do not correspond to any concept in our lexicon. Superposition, wherein neurons encode multiple unrelated features simultaneously, further complicates the picture. Models seem to develop both human-like abstractions and genuinely alien ones, interleaved in ways we are only beginning to disentangle.

The honest reading of the current evidence is that convergence is partial and domain-dependent. High-level semantic concepts appear to converge more than low-level perceptual features. Whether this pattern will hold or reverse as capabilities scale further is one of the most consequential open questions in AI research today.

Takeaway

Interpretability findings hint that artificial minds partially share our conceptual furniture—but they also arrange alien objects in rooms we have not yet learned to enter.

The Alignment Wager

For AI alignment, the stakes of this question are difficult to overstate. If human and machine concepts converge, then training an AI on human feedback becomes a project of refining a shared vocabulary rather than translating across an unbridgeable chasm. Value learning becomes tractable because the AI's internal notion of, say, harm or honesty would plausibly resemble our own, differing in emphasis rather than in kind.

This is the optimistic reading. Shared abstractions would make interpretability meaningful, oversight coherent, and course-correction feasible. When we ask a model to explain its reasoning, its self-report would refer to concepts we recognize, because those concepts are the natural currency of any mind that has modeled our world with fidelity.

The pessimistic reading is more unsettling. Partial convergence may be worse than either extreme, because it creates the illusion of shared understanding while concealing critical divergences. An AI whose concept of human wellbeing overlaps with ours by ninety-five percent may pursue goals that appear aligned in every test scenario yet catastrophically diverge in edge cases we failed to probe. Surface agreement can mask ontological drift.

Stuart Russell's articulation of the control problem takes on new dimensions here. If we cannot verify that an AI's abstractions match ours in the ways that matter, we cannot verify that its objectives, however elegantly specified, will remain safe under distribution shift. The natural abstractions hypothesis, if true, is a gift to alignment researchers. If false, or true only partially, it may be a trap that lulls us into false confidence.

The deepest implication may be epistemic rather than technical. We are, for the first time, in a position to empirically investigate whether our concepts are provincial or universal—whether the human mind was tracking reality or merely painting it. The answer, whatever it turns out to be, will reshape not only how we build machines but how we understand ourselves.

Takeaway

Partial conceptual overlap between minds is the most dangerous kind, because it produces fluent conversation without guaranteeing genuine agreement about the world.

The natural abstractions hypothesis is ultimately a bet about the depth of the correspondence between mind and world. It suggests that intelligence, wherever it arises and whatever substrate hosts it, is drawn toward the same conceptual landmarks by the informational geography of reality itself.

The evidence to date is suggestive rather than decisive. We see convergence and divergence, familiar concepts and alien ones, in the same networks. What we do not yet possess is a rigorous map of which abstractions are natural and which are contingent—a map that alignment may ultimately require.

Perhaps the most valuable stance is disciplined hope: to work as though shared abstractions are achievable while auditing relentlessly for the places they fail. The alternative—assuming either total alignment or total alienness—forecloses the careful empirical work this moment demands. Reality's joints, if they exist, will not reveal themselves to those who stop looking.