How does the brain encode a concept? For decades, this question has oscillated between two poles. On one extreme sits the localist view, caricatured by the infamous grandmother cell: a single neuron dedicated to a single concept. On the other lies the fully distributed view, where every concept is a diffuse pattern across millions of neurons, no single unit meaningful in isolation.

Both extremes are theoretically elegant, and both are empirically wrong. The nervous system, as it turns out, occupies a strange middle territory—one that neither classical connectionism nor extreme localism anticipated. Recordings from human medial temporal lobe reveal neurons of astonishing selectivity. Simultaneously, population-level analyses show that meaning is smeared across ensembles. These findings are not contradictions to be resolved but clues to a deeper computational principle.

The brain, it appears, tunes its coding scheme along a continuum of sparsity, optimising representations for the computational demands of each region. Sensory peripheries favour dense codes to maximise information throughput; associative regions favour sparse codes to enable rapid associative learning and pattern separation. Understanding this gradient—why it exists, how it emerges, and what it implies—reframes some of the most fundamental questions in theoretical neuroscience, from memory formation to the neural correlates of conscious content.

Concept Cells and the Ghost of the Grandmother

In 2005, Quiroga and colleagues published a finding that unsettled the distributed-representation orthodoxy. Recording from single neurons in the human medial temporal lobe of epilepsy patients, they identified units that responded selectively to specific individuals—one famously to Jennifer Aniston, another to Halle Berry—across photographs, drawings, and even the written name of the person.

These concept cells exhibit invariance that classical feature-detector models cannot explain. A single neuron firing to both an image and a printed name has abstracted away modality entirely, encoding something closer to semantic identity than to sensory features. This is not the grandmother cell of Lettvin's original thought experiment, but it is uncomfortably close.

Yet the resemblance is superficial. Concept cells are not solitary. Estimates suggest that a given concept engages perhaps 40 to 150 such neurons within a population of roughly a million in the hippocampus and neighbouring cortex. This is sparse, but not local. The redundancy provides robustness against neural death and noise, while the sparsity enables the rapid formation of new associations without catastrophic interference.

Theoretically, this coding regime resembles what computational models call a sparse expansion: a projection from denser sensory representations into a higher-dimensional space where only a small fraction of units are active for any input. The mathematics of such expansions—studied in models from Marr's cerebellum through modern hippocampal theories—show they are ideal substrates for pattern separation and associative memory.

The concept cell, then, is not evidence for localism. It is evidence that the brain has discovered a specific computational sweet spot, one where selectivity is high enough for efficient indexing but distributed enough for graceful degradation and combinatorial flexibility.

Takeaway

Selectivity is not the same as locality. A neuron can be remarkably specific in what it represents while still participating in a distributed code—the question is not whether representations are localised, but how sparse they are.

The Information-Theoretic Case for Distribution

Why would evolution favour distributed coding at all? The answer emerges from information theory. Consider a population of N binary neurons. A purely localist code represents at most N concepts. A fully distributed binary code can represent 2N patterns—an astronomically larger capacity.

But capacity alone does not dictate optimality. Fully dense codes—where roughly half the population is active for any input—suffer catastrophic interference in associative learning. Any two overlapping patterns interfere; every new memory disturbs the old. The Hopfield network, that classical crystallisation of distributed memory, has a storage capacity of only ~0.14N patterns before retrieval fails.

Sparse distributed codes resolve this tension. With activation probability p ≪ 0.5, the expected overlap between random patterns falls dramatically, and associative capacity scales as roughly N²/(p log(1/p)). This is orders of magnitude beyond the dense regime. The mathematical convenience aligns with a biological constraint: metabolic cost. Action potentials are expensive, and sparsity minimises energetic expenditure per bit of information transmitted.

There is also a decodability advantage. Downstream neurons performing linear readouts can extract information from sparse codes with simple synaptic weightings, whereas dense codes often require nonlinear separation. The brain's abundance of Hebbian-like plasticity rules is well-matched to sparse representations, where coincident activity is rare and therefore meaningful.

Distribution, in short, is not a philosophical commitment but a computational necessity. The question is not whether to distribute but how sparsely—and that answer depends on what the region is computing.

Takeaway

Sparsity is the negotiated peace between representational capacity and interference. The brain does not maximise information density; it optimises the trade-off between what can be stored and what can be retrieved cleanly.

Task-Dependent Sparsity Across the Cortical Hierarchy

The most compelling evidence against a one-size-fits-all coding scheme comes from comparing sparsity across brain regions. Primary visual cortex operates with roughly 10-20% of neurons active for a given stimulus. In inferotemporal cortex, that figure drops to perhaps 1-5%. In the hippocampus and entorhinal cortex, active fractions can fall below 1%. This is a systematic gradient, not noise.

The gradient tracks computational function. Early sensory areas must transmit high-bandwidth information about the world—edges, orientations, colours—and benefit from denser codes that carry maximal Shannon information per spike. Higher cortical areas, tasked with categorisation, memory, and abstraction, benefit from sparser codes that enable pattern separation and rapid learning.

This is not merely descriptive. Theoretical models by Olshausen, Field, and successors demonstrated that sparse coding objectives, when applied to natural image statistics, spontaneously produce receptive fields resembling those of V1 simple cells. The brain appears to have discovered, through evolution and development, the coding scheme that best matches the statistical structure of its inputs at each processing stage.

Moreover, sparsity is not static. Attention, task demands, and behavioural state modulate activation fractions dynamically. Under focused attention, cortical responses sharpen and populations become sparser; under generalised arousal, they broaden. The brain does not commit to a fixed coding scheme—it tunes sparsity in real time as computational demands shift.

This tunability implies a deeper principle: coding sparsity is itself a controlled variable, adjusted by neuromodulatory and inhibitory circuits to match the computational regime the organism currently requires.

Takeaway

There is no universal neural code. The brain uses a spectrum of coding schemes, and sparsity is calibrated—both anatomically and dynamically—to the computational task at hand.

The dichotomy between localist and distributed representations was always a false one, a residue of theoretical convenience rather than empirical reality. The brain occupies a rich intermediate landscape, employing sparse distributed codes whose density is tuned to the computational demands of each region and moment.

This reframing has implications far beyond coding theory. Memory capacity, generalisation, one-shot learning, and even the phenomenology of conscious perception all depend on the sparsity regime of the underlying substrate. Theories of consciousness that posit specific neural correlates must reckon with the fact that those correlates are neither single cells nor undifferentiated populations, but structured ensembles with characteristic geometries.

Perhaps the deepest lesson is methodological. Progress in theoretical neuroscience will not come from choosing between elegant extremes but from characterising the continuum along which the brain actually operates—and understanding why evolution parked each region where it did.