Government budgets allocate trillions of dollars annually across programs whose actual impacts remain, in many cases, empirically undetermined. This constitutes a fundamental information failure in public finance: resources flow according to political salience, administrative inertia, and stakeholder pressure rather than measured social returns. Optimal fiscal policy requires closing this gap.

The methodological revolution in applied microeconomics—the credibility revolution associated with randomized experiments, regression discontinuity designs, and quasi-experimental methods—has produced tools capable of generating rigorous causal estimates of program effects. Yet the translation of these methods into routine government practice remains incomplete, and the translation of findings into budget decisions is weaker still.

This article develops a framework for integrating credible causal inference into public resource allocation. We examine three interconnected challenges: selecting identification strategies appropriate to policy settings, combining impact estimates with cost data to produce decision-relevant metrics, and diagnosing the institutional pathologies that prevent evaluation findings from shaping budgets. The stakes extend beyond individual programs. A public sector that systematically learns from its own interventions can approach something resembling optimal expenditure design; one that does not is condemned to allocate by intuition, ideology, and inertia. Building the institutional infrastructure for evidence-based budgeting is itself a public good with substantial fiscal returns.

Identification Strategies for Government Programs

The core challenge of program evaluation is distinguishing the causal effect of an intervention from the counterfactual—what would have occurred absent the program. Naive comparisons of participants and non-participants routinely conflate treatment effects with selection effects, producing biased estimates that can mislead resource allocation by orders of magnitude.

Randomized controlled trials, when feasible, provide the cleanest identification. By assigning eligibility or intensity randomly across otherwise identical units, RCTs eliminate selection bias by construction. The Oregon Health Insurance Experiment, the Moving to Opportunity demonstration, and numerous randomized evaluations of active labor market programs illustrate how experimental variation can resolve otherwise intractable causal questions about consequential policies.

Where randomization is politically or ethically infeasible, quasi-experimental methods exploit variation generated by policy rules themselves. Regression discontinuity designs leverage sharp eligibility thresholds—income cutoffs for transfers, test-score cutoffs for scholarships, age cutoffs for benefits—to compare units on either side of arbitrary boundaries. Difference-in-differences designs exploit staggered policy adoption across jurisdictions. Instrumental variables approaches use exogenous shocks to program participation to recover local average treatment effects.

Each method carries assumptions that must be defended, not assumed. Parallel trends must be examined, not asserted. Exclusion restrictions must be argued substantively. Bandwidth choices in RD designs must be tested for sensitivity. The credibility of any estimate depends on the transparency with which its identifying assumptions are stated and probed.

Government evaluation infrastructure should be designed to generate exogenous variation deliberately: phased rollouts randomized across sites, lottery-based allocation when demand exceeds capacity, and administrative rules with sharp discontinuities. Treating implementation choices as opportunities for learning transforms routine program administration into a continuous evidence-generation apparatus.

Takeaway

Every policy design decision is also a research design decision. Governments that build randomization and sharp rules into implementation generate causal evidence as a byproduct of doing their work.

Integrating Impact Estimates with Cost Data

A credibly identified treatment effect is a necessary but insufficient input to budget decisions. Programs must be evaluated on their social returns per dollar spent, not merely on whether their effects are statistically distinguishable from zero. The relevant question is comparative: given a fixed budget envelope, which allocation across programs maximizes social welfare?

The Marginal Value of Public Funds framework, developed by Hendren and Sprung-Keyser, formalizes this comparison. For each policy, one computes the ratio of beneficiaries' willingness to pay for the intervention to its net government cost, accounting for behavioral responses and fiscal externalities. Policies with MVPF above unity generate more social value than the resources they consume; those below unity destroy value at the margin.

This framework reveals that headline impact estimates can be deeply misleading. A program with modest measured effects may nonetheless dominate alternatives if its costs are low or if it generates substantial fiscal savings elsewhere—reduced Medicaid utilization, higher future tax revenues, lower incarceration expenditures. Conversely, programs with impressive raw effects may perform poorly once full costs, including deadweight losses from financing, are accounted for.

Rigorous cost-benefit integration requires long-run outcome tracking, since many high-return interventions—early childhood programs, preventive health investments, human capital subsidies—generate benefits over decades. Administrative data linkages that follow cohorts across tax, health, education, and criminal justice systems are essential infrastructure. Without them, evaluations systematically undervalue programs whose returns materialize slowly.

Distributional weighting introduces further complexity. A dollar of benefit to a low-income household plausibly generates more social welfare than a dollar to a high-income household, but the choice of weights is normative and should be made transparent. Presenting MVPF estimates under alternative distributional assumptions allows decision-makers to see how conclusions depend on value judgments they must ultimately own.

Takeaway

Impact without cost is a partial truth; cost without impact is bureaucratic accounting. Only their integration—expressed as social value per public dollar—provides the metric optimal budgeting requires.

Institutional Barriers to Evidence Use

The technical production of credible evaluations does not automatically translate into their use. Decades of investment in program evaluation have produced substantial evidence bases whose influence on actual budget allocations remains disappointingly limited. Understanding why requires examining the political economy of evidence use, not merely the epistemology of evidence production.

Legislatures allocate resources through processes shaped by concentrated beneficiary interests, agency advocacy, and electoral incentives. Evaluation findings enter this process as one input among many, and rarely the most politically weighty. Programs with strong constituencies survive evidence of ineffectiveness; programs without champions can be terminated regardless of demonstrated returns. The Congressional Budget Office and analogous institutions provide some counterweight, but their mandates typically emphasize fiscal scoring over impact assessment.

Agency incentives compound the problem. Bureaus that commission evaluations of their own programs face asymmetric risks: positive findings offer modest reputational gains, while negative findings threaten budgets and careers. This selection dynamic biases the evaluation portfolio toward safe questions and away from consequential ones. Independent evaluation authorities, insulated from operational agencies, partially address this pathology.

Institutional designs that improve evidence use share common features. Statutory requirements for evaluation, as in the Foundations for Evidence-Based Policymaking Act, create baseline expectations. Sunset provisions and reauthorization requirements force periodic reconsideration. Evidence clearinghouses that systematically synthesize findings across studies reduce the cognitive burden on decision-makers. Learning agendas that align evaluation priorities with budget cycles ensure findings arrive when they can be used.

The deeper reform is cultural. Governments that treat evaluation as a compliance exercise produce compliance-quality evidence. Those that treat it as core to their mission—embedding chief evaluation officers with real authority, protecting evaluators from operational pressure, and creating career paths that reward analytical rigor—generate evidence that actually shapes decisions. The institutional architecture of learning is itself a policy choice.

Takeaway

Evidence does not use itself. Whether rigorous findings shape budgets depends less on their methodological quality than on institutions designed to translate research into resource allocation.

Program evaluation is not an academic ornament to public administration but a core input to optimal fiscal policy. The methodological tools now exist to estimate program effects credibly; the analytical frameworks exist to translate those estimates into welfare-relevant metrics. What remains underdeveloped is the institutional machinery that channels evidence into budget decisions.

Building this machinery requires simultaneous investment on three fronts: designing programs to generate causal variation, linking administrative data to enable long-run cost-benefit analysis, and constructing evaluation institutions insulated from the operational pressures that distort evidence production. None of these investments yields immediate political returns, but their cumulative effect is a public sector capable of learning from its own actions.

The alternative—allocating trillions annually by intuition, tradition, and interest-group pressure—represents a substantial ongoing welfare loss. A government that takes its own effectiveness seriously must take evidence seriously, not as rhetoric but as institutional practice.