The randomized controlled trial revolution has transformed development economics, delivering unprecedented rigor to questions about what works in reducing poverty. Yet a systematic blind spot persists in our evaluation architecture: we routinely measure impact at one, two, or three years post-intervention and treat those estimates as definitive verdicts on program effectiveness. This truncated time horizon may be quietly distorting our understanding of development itself.
Consider the mismatch. The theories of change underlying most development interventions—improved cognitive development, altered household investment patterns, disrupted poverty traps—operate on timescales of decades. Yet the empirical evidence we generate to test these theories rarely extends beyond a graduate student's dissertation window. We are, in effect, running short-exposure film to capture long-exposure phenomena.
The consequences are more than academic. Programs are scaled up on the basis of promising short-term effects that later dissipate. Interventions with modest immediate impact but transformative long-run effects are prematurely abandoned. Cost-effectiveness calculations that shape billions in development spending rest on truncated impact estimates. If we are serious about evidence-based development, we must reckon with the temporal architecture of the evidence itself—and confront why longer horizons remain the exception rather than the norm.
Fade-Out Patterns: When Promising Effects Disappear
The fade-out phenomenon has haunted development evaluation for decades, most famously in the education literature. Preschool interventions that generate substantial short-term gains in test scores often show attenuated or null effects on academic outcomes measured five to ten years later. The Head Start Impact Study documented this convergence pattern with unusual rigor, finding that initial cognitive advantages had largely dissipated by third grade.
Similar patterns appear across sectors. Teacher training programs in Kenya evaluated by Duflo, Dupas, and Kremer showed learning gains that eroded once implementing partners exited. Microfinance evaluations documenting business expansion at eighteen months revealed muted effects on consumption and poverty at longer horizons. Deworming's contested long-run effects illustrate how contested this terrain becomes when we actually attempt to measure it.
The mechanisms driving fade-out matter enormously for program design. Sometimes control groups catch up through spillovers or subsequent interventions—a rising tide that erodes the treatment differential without invalidating the original impact. Other times, treatment effects genuinely dissipate because complementary inputs are missing, human capital depreciates without reinforcement, or general equilibrium adjustments neutralize partial-equilibrium gains.
This distinction is crucial but rarely investigated. A program showing fade-out because everyone eventually receives similar services tells a very different policy story than one where treatment participants regress to counterfactual outcomes. Yet without carefully designed long-term follow-ups that measure mechanisms alongside outcomes, we cannot distinguish these scenarios.
The uncomfortable implication is that many programs celebrated for strong short-term effects may be delivering less durable value than we assume. Publication bias compounds the problem: fade-out results are harder to publish, less career-enhancing to pursue, and more damaging to the programmatic reputations of those who funded the original work.
TakeawayA treatment effect measured at two years is not an estimate of program impact—it is an estimate of program impact at two years. Treating the former as the latter has cost the development community dearly in misallocated resources and misplaced confidence.
Sleeper Effects: When Impact Emerges Years Later
The mirror image of fade-out is equally consequential and even more neglected: interventions whose true impact emerges only after substantial delay. The Perry Preschool and Abecedarian studies, followed for decades, revealed that early childhood investments producing modest cognitive gains generated substantial effects on adult earnings, incarceration, and health—effects that would have been entirely invisible in any conventional evaluation window.
The theoretical rationale is straightforward. Human capital investments compound over time. Behavioral and social skills shape trajectories through education, labor markets, and family formation—each transition amplifying initial advantages. Yet these dynamic multiplier effects only manifest across life-cycle transitions that occur far outside typical evaluation timelines.
Chetty and coauthors' work on neighborhood effects and teacher quality demonstrates how administrative data linkage can retrospectively recover such long-run impacts. Their finding that kindergarten classroom quality predicts adult earnings—despite modest effects on test scores that fade by later grades—should permanently reshape how we interpret null medium-term results.
The policy implications cut both ways. Some interventions we have abandoned as ineffective may have been quietly generating substantial long-run value. Others we celebrate for measurable short-term gains may be delivering less than their long-run competitors. Cost-effectiveness rankings built on medium-term outcomes may systematically misorder our programmatic priorities.
Sleeper effects also raise uncomfortable questions about intergenerational transmission. If maternal health interventions today alter children's outcomes twenty years hence, and their children's outcomes forty years hence, the true benefit-cost ratio of many programs may be dramatically understated by any tractable evaluation horizon.
TakeawayThe absence of measurable short-term impact is not evidence of absent long-term impact. In development economics, some of the most valuable interventions may look unimpressive precisely when we choose to measure them.
Practical Barriers: Why Long-Term Follow-Up Remains Rare
The structural obstacles to long-term evaluation are considerable and reinforce one another. Research funding cycles typically span three to five years, aligned poorly with the decade-plus horizons required to observe compounding effects. Grant renewal depends on published output, which depends on evaluation results, which depends on early endline measurement. The incentive gradient runs sharply against patience.
Researcher career incentives compound the problem. Junior scholars cannot wait fifteen years for tenure-track publications. Long-term follow-ups typically require substantial upfront investment in tracking systems, biometric identification, or administrative data linkage—infrastructure that generates no immediate publications and whose payoff accrues largely to future researchers who inherit the cohort.
Attrition presents genuine methodological challenges. In highly mobile populations, tracking respondents over a decade may cost more than the original evaluation. Selective attrition threatens internal validity in ways that cross-sectional endlines do not face. These are real constraints, not merely excuses—but they are constraints that can be systematically addressed through investment.
Several solutions merit serious consideration. Dedicated long-term evaluation funds, insulated from typical grant cycles, could underwrite decade-plus follow-ups on high-value interventions. Government administrative data infrastructure—tax records, health registries, education databases—dramatically reduces tracking costs when accessible to researchers. Consortium models allow multiple studies to share tracking infrastructure across a common cohort.
The J-PAL Long-Term Follow-Up Fund represents one promising institutional response, explicitly funding revisits to previously evaluated interventions. Similar dedicated resources at bilateral donors and foundations could substantially shift the temporal composition of the evidence base. What we measure reflects what we fund; changing the funding architecture is the necessary precondition for changing the evidence.
TakeawayThe temporal shortsightedness of development evidence is not an intellectual failure but an institutional one. Fixing it requires deliberately restructuring incentives, not merely exhorting researchers to think longer-term.
Rigorous evaluation was never meant to deliver instant verdicts on complex social interventions. The credibility revolution in development economics has given us powerful tools for causal identification; we now need matching investment in temporal reach. Without it, we risk building a policy edifice on empirical foundations that measure the wrong horizon of the phenomena we care about.
This does not diminish the value of short and medium-term evaluations. Such studies remain essential for detecting implementation failures, refining program design, and generating early evidence for iterative learning. The point is not to replace them but to complement them systematically with the longer-horizon follow-ups that alone can validate—or overturn—their conclusions.
Development is, by its nature, a long-run process. Our evidence architecture should reflect that reality rather than fight against it. The interventions that ultimately transform lives may look quite different from those that maximize measurable impact within a graduate student's timeline.