Two decades ago, development evaluation was a small, marginal activity. A handful of economists ran studies. Programs were largely judged by inputs delivered and outputs counted. Impact was assumed, not measured.
Today, evaluation is a multi-billion dollar industry. Every major donor has an evaluation department. Consulting firms compete for contracts worth tens of millions. Randomized controlled trials have won Nobel Prizes and colonized development discourse. The shift represents a genuine intellectual victory for evidence-based thinking.
But industries develop their own logic. When evaluation becomes a career, a business line, and a compliance requirement, its incentives change. This article examines whether the evaluation boom has actually improved aid effectiveness—or whether it has produced a sophisticated new form of theater that satisfies donors without transforming practice.
The Evaluation Explosion
The numbers tell a striking story. In 2000, the World Bank's Independent Evaluation Group had a modest budget and produced a limited portfolio of studies. By 2020, global spending on development evaluation exceeded an estimated $2 billion annually, with major bilateral donors like USAID, DFID, and GIZ each spending hundreds of millions on evaluation activities.
The Abdul Latif Jameel Poverty Action Lab has run over 1,000 randomized evaluations across 90 countries. Innovations for Poverty Action, 3ie, and dozens of academic centers have institutionalized rigorous impact assessment. Every major foundation now expects logframes, theories of change, and measurable indicators as baseline requirements.
This growth reflects genuine progress. The credibility revolution in economics brought experimental methods that could actually identify what works. Cash transfer research, deworming debates, and microfinance studies produced findings that changed policy. Donors who once funded projects on faith increasingly demand evidence.
Yet the volume of evaluation has grown far faster than its influence on practice. Studies pile up in institutional repositories. Meta-analyses reveal that most evaluations are never cited by anyone making programming decisions. The infrastructure has scaled dramatically; the learning has not kept pace.
TakeawayAn activity can expand enormously while its original purpose atrophies. Measuring more is not the same as learning more.
Quality and Independence Questions
Evaluation is now a market, and markets have incentives. Consulting firms bidding for evaluation contracts know that findings deeply critical of a program rarely lead to follow-on work. Implementing organizations often select and pay their own evaluators. The evaluator who finds a program transformative gets rehired; the one who finds it ineffective may not.
Research by Eva Vivalt and colleagues has documented publication bias and heterogeneity problems even in academic impact evaluations. Effect sizes shrink dramatically when studies are replicated in new contexts. Yet program designs continue to be justified by pointing to a handful of positive studies, often from very different settings.
Independence is compromised in subtler ways too. Evaluators embedded in the development ecosystem share the assumptions, vocabulary, and career interests of the organizations they assess. Truly damaging findings—that a flagship program does nothing, or actively harms recipients—face social and financial headwinds before they even reach a draft report.
The result is a peculiar equilibrium: methodologically sophisticated studies producing conclusions that rarely upset donor priorities. Rigor at the technical level coexists with predictability at the conclusion level. The word 'promising' does enormous work in obscuring what the data actually shows.
TakeawayIndependence is not just about who signs the paycheck. It is about whether critical findings can survive the social and institutional pressures they will inevitably face.
From Learning to Accountability
Evaluation has always served two purposes uneasily. Learning asks: what should we do differently? Accountability asks: did we do what we promised? These require different methods, different audiences, and often different findings—but they now share the same reports.
Accountability pressures dominate. Donors need evidence for parliaments and taxpayers. Implementing organizations need proof of value for renewal. Recipients need to demonstrate compliance for continued funding. In this climate, evaluation becomes a document produced to justify decisions already made, not a tool to inform future ones.
The learning function suffers accordingly. Genuine learning requires admitting failure, pursuing surprising findings, and revising theories of change. But an accountability-driven evaluation cannot admit failure without triggering funding consequences. Programs are quietly redefined so their original goals become their measured outcomes. Success is engineered backward from whatever the data happens to show.
The most useful evaluations in development history—the deworming debates, the microfinance reckoning, the graduation program findings—required years of contested analysis and willingness to disturb established narratives. They emerged despite the accountability system, not because of it. Institutional pressures now push in the opposite direction.
TakeawayWhen one tool must serve two masters, it usually ends up serving the more powerful one. Development evaluation has quietly chosen accountability over learning.
The evaluation industry has delivered real gains. We know more about what works in development than we did twenty years ago. Cash transfers, chlorinated water, and targeted interventions have accumulated genuine evidence bases.
But industrialized evaluation has also produced industrialized rationalization. Sophisticated methods now regularly validate programs whose theoretical foundations remain shaky, whose contextual applicability is unproven, and whose actual effects on poverty are marginal.
The path forward requires separating learning from accountability, protecting evaluators who deliver uncomfortable findings, and treating negative results as valuable knowledge. Otherwise, we have built an expensive apparatus for producing the appearance of evidence rather than the substance of it.