In 1969, when the Apollo 11 astronauts placed a retroreflector on the lunar surface, physicists on Earth began firing laser pulses at the Moon to measure its distance with unprecedented precision. But how did they know their measurements were correct? Against what standard could you possibly calibrate a device measuring something no one had ever measured before?
This question sits at the heart of experimental science, though it rarely receives the attention it deserves. We tend to imagine instruments as neutral windows onto reality—thermometers reveal temperature, telescopes reveal distant galaxies, particle detectors reveal subatomic events. Yet every instrument embodies theoretical assumptions, engineering compromises, and interpretive conventions. Trusting an instrument is never a simple act of perception; it is a sophisticated cognitive and social achievement.
The philosopher Harry Collins and historian Peter Galison have spent decades illuminating how experimental physicists navigate this territory. What emerges from their work is a picture of science as a discipline that has developed remarkable methods for building trust in its tools—but also one that occasionally stumbles into deep uncertainty about whether what it sees is real. Understanding how scientists come to believe their instruments reveals something profound about the nature of scientific knowledge itself: that it is built not on unmediated observation but on carefully constructed chains of inference and community consensus.
Calibration Networks
Every measurement in science ultimately traces back through a genealogy of prior measurements. A laboratory scale is calibrated against reference weights, which were calibrated against national standards, which were calibrated against international prototypes maintained at the Bureau International des Poids et Mesures in Sèvres. This is a calibration network—a web of interconnected instruments whose collective reliability exceeds what any single device could achieve alone.
The genius of calibration networks lies in their redundancy. When multiple independent instruments, built on different principles, converge on the same measurement, we gain confidence that we are measuring something real rather than artifacts of any particular technique. Astronomers measuring cosmic distances use parallax, standard candles, and redshift methods that overlap in certain regimes, allowing each to check the others.
But calibration is never purely mechanical. It requires judgment about which reference standards are appropriate, which comparison conditions matter, and which discrepancies are meaningful versus which fall within acceptable tolerance. New fields face a particular challenge: without established standards, they must bootstrap their measurement systems, often by comparing novel instruments against theoretical predictions or against each other in a delicate coordination.
The historian Simon Schaffer has shown how early electrical measurements in the nineteenth century required physical shipments of standard resistors and capacitors between laboratories, along with skilled technicians who could demonstrate proper technique. Trust traveled with objects and bodies, not merely with numbers on paper.
This social dimension persists today. When a laboratory claims a new measurement, its acceptance depends on whether the broader community can situate that measurement within existing calibration networks—or whether the claim requires extending those networks in ways others can verify.
TakeawayNo instrument stands alone. Every trusted measurement is a node in a vast network of cross-checks, and confidence emerges from the coherence of the whole rather than the authority of any single device.
The Experimenter's Regress
Harry Collins identified a troubling circularity in experimental science that he called the experimenter's regress. To know whether an experiment has been performed correctly, we need to know whether it produces the correct result. But to know what the correct result is, we need experiments performed correctly. This regress becomes acute precisely when it matters most: at the frontiers of research, where no prior consensus exists.
Collins studied the search for gravitational waves in the 1970s, when Joseph Weber claimed to detect them using aluminum bar detectors. Other physicists built similar detectors and saw nothing. Was Weber's apparatus flawed, or were the others insufficiently sensitive? Each side could accuse the other of running their experiments incorrectly, and there was no independent standard to adjudicate the dispute.
The regress is only broken through social processes—through negotiations, replications, tacit demonstrations of expertise, and eventually the emergence of community consensus. This is not a failure of science but a recognition of how knowledge is actually produced at the edge of the known. Consensus does not follow automatically from data; it must be constructed through argument and practice.
What makes the experimenter's regress philosophically interesting is that it reveals the impossibility of purely algorithmic science. There is no rule book that tells you when your apparatus is working correctly in the absence of trusted results to compare against. Scientific judgment is required, and that judgment is developed through immersion in a research community.
The regress also explains why controversial results often take decades to resolve. What looks like stubbornness or bias from outside is often a legitimate uncertainty about which experimental configuration to trust when the phenomenon itself remains contested.
TakeawayAt the frontiers of knowledge, data alone cannot settle disputes about what data means. Science advances not by pure logic but by the patient construction of shared judgment within expert communities.
Instrumental Expertise and Tacit Knowledge
Michael Polanyi observed that we know more than we can tell. Nowhere is this truer than in laboratory practice, where experienced researchers develop a feel for their instruments that resists complete articulation. A skilled electron microscopist can sense when the beam is drifting; a veteran chemist can tell when a reaction is proceeding correctly by subtle changes in color or smell. This is tacit knowledge—embodied expertise that cannot be fully captured in manuals.
The importance of tacit knowledge becomes visible when techniques fail to transfer between laboratories. Collins documented how researchers trying to build TEA lasers in the 1970s repeatedly failed until they visited a laboratory where the technique worked, absorbing through direct observation the countless small adjustments that written protocols omitted. Some knowledge must pass through bodies, not just papers.
Instrumental expertise involves knowing when to trust a reading and when to distrust it—recognizing the signature of a genuine signal versus artifact, understanding which quirks of the apparatus matter and which do not. This diagnostic skill develops through prolonged engagement, often accompanied by failures that teach lessons no textbook could convey.
The rise of automated instruments and standardized commercial equipment has not eliminated tacit knowledge; it has relocated it. Modern researchers must know when to trust the black box and when to question it, when default settings suffice and when custom configurations are required. The skill has shifted from operating the instrument to interrogating it.
This has profound implications for how new scientific fields develop. Establishing a new research area requires more than building new instruments—it requires cultivating a generation of practitioners who possess the embodied know-how to use them well. The transmission of expertise remains stubbornly artisanal, even in our data-driven age.
TakeawayScientific mastery is not merely intellectual but bodily and communal. The most important knowledge in any laboratory often lives in hands and habits, transmitted through apprenticeship rather than instruction.
The question of how scientists come to trust their instruments turns out to be one of the deepest questions in the philosophy of science. It reveals that scientific objectivity is not a matter of stepping outside human judgment but of cultivating a particular kind of disciplined, communal judgment.
This has practical implications for anyone working at research frontiers. New fields must invest not only in new equipment but in the slow work of building calibration networks, developing shared diagnostic vocabularies, and training practitioners whose embodied expertise can distinguish signal from noise. These investments look inefficient from outside but are prerequisites for reliable knowledge.
Perhaps most importantly, understanding calibration and trust should make us both more humble about scientific claims and more appreciative of them. When science works—when instruments reveal genuine features of the world—it is because generations of careful practitioners have woven together a fabric of trust that no single measurement could sustain alone.