Performance budgeting represents one of the most ambitious institutional experiments in modern public finance: the attempt to systematically align resource allocation with measurable outcomes. The theoretical appeal is straightforward. If governments could credibly link funding to results, they would generate powerful incentives for efficiency, transparency, and evidence-based policymaking. Yet after four decades of experimentation across OECD economies, the empirical record remains stubbornly mixed.

The core challenge is not technical but structural. Performance information exists at the intersection of principal-agent problems, incomplete contracts, and multi-task incentive design. Agencies deliver services with heterogeneous outputs, contested objectives, and outcomes shaped by exogenous factors. Simply measuring more, or tying budgets more tightly to metrics, does not resolve these frictions—it often amplifies them through gaming, goal displacement, and measurement fatigue.

This article develops a design framework for performance budgeting grounded in mechanism design and information economics. We examine three interlocking problems: how to construct meaningful indicator systems, how to integrate performance data into allocation decisions along a spectrum from informational to formulaic, and how to structure incentives that minimize behavioral distortion. The goal is not to advocate for or against performance budgeting in the abstract, but to identify the institutional configurations under which it plausibly enhances welfare, and those under which it becomes an expensive theater of accountability.

Information Architecture: Designing Indicators That Signal Rather Than Distort

The foundational task in performance budgeting is constructing an indicator system that carries genuine informational content about program effectiveness. This is substantially harder than practitioners typically acknowledge. Holmström's classic multi-task model demonstrates that when agents produce multiple outputs of varying measurability, high-powered incentives on measurable dimensions predictably crowd out unmeasured but socially valuable activities. Indicator design is therefore not a technical exercise in metric selection but an exercise in incomplete contract theory.

A well-designed indicator system distinguishes rigorously between inputs, outputs, intermediate outcomes, and final outcomes. Input and output measures are typically verifiable but weakly correlated with social welfare. Outcome measures track welfare more directly but suffer from attribution problems, time lags, and exogenous shocks that add noise to any inference about program performance. The optimal indicator portfolio balances these tradeoffs, often combining a small set of headline outcome indicators with process measures that capture managerial effort.

Validation is equally critical. Self-reported administrative data suffers from well-documented distortions: recategorization of cases, selective reporting, and outright fabrication under high-stakes conditions. Robust systems triangulate administrative data with independent audits, survey-based measures, and quasi-experimental impact evaluations. The Government Performance and Results Modernization Act framework, for instance, remains vulnerable precisely because it lacks systematic external validation of agency-reported metrics.

The number of indicators matters as much as their content. Cognitive and organizational research consistently shows that when agencies track dozens of indicators, none receive meaningful attention. Effective systems typically converge on five to ten priority indicators per program, with deeper diagnostic measures used episodically for evaluation rather than continuous monitoring. Parsimony is a design virtue, not a limitation.

Finally, indicator systems must accommodate learning. Static metrics locked in through statute or regulation become obsolete as programs evolve, technologies change, and populations shift. Institutional mechanisms for periodic indicator revision—ideally insulated from short-term political pressure but subject to democratic oversight—are essential features of any durable performance architecture.

Takeaway

Every indicator is an incomplete contract. The question is not whether measurement distorts behavior, but whether the distortions it creates are less costly than the informational asymmetries it resolves.

Budgetary Integration: The Spectrum from Information to Formula

Once performance information exists, the question becomes how tightly to couple it with resource allocation. This design choice spans a spectrum, and the optimal position depends critically on the informational quality of the metrics and the political economy of the allocation process. Understanding this spectrum clarifies why identical performance systems produce radically different outcomes across jurisdictions.

At the loosest coupling, performance-informed budgeting simply provides decision-makers with performance data alongside traditional budget requests. This approach, adopted in various forms across Nordic countries, treats performance information as one input among many into a fundamentally deliberative process. It preserves flexibility, allows contextual judgment, and minimizes gaming incentives—but it also risks performance data being ignored entirely when it conflicts with political priorities.

Intermediate models introduce structured performance dialogues, performance agreements, or conditional funding tranches. The UK's Public Service Agreements and various OECD spending review frameworks exemplify this approach. Performance data shapes negotiations without mechanically determining outcomes. Empirical evaluation suggests these models can meaningfully influence resource allocation when supported by strong central capacity, but they demand substantial analytical infrastructure and executive attention.

At the tightest coupling, formula-based performance funding directly ties allocations to measured outputs or outcomes. Higher education funding formulas that reward completions, healthcare capitation adjusted for outcomes, and results-based aid disbursement operate on this principle. The efficiency gains from high-powered incentives are real but come at substantial cost: intense gaming pressure, risk aversion in program design, and cream-skimming of easier-to-serve populations.

The optimal position on this spectrum is not universal. It depends on measurement quality, the observability of exogenous shocks, the risk preferences of program managers, and the political salience of the outcomes in question. Programs with well-measured, attributable outcomes and low political stakes can sustain tighter coupling. Programs with noisy metrics and high political salience typically require looser integration to avoid destructive gaming.

Takeaway

Tight coupling between metrics and money creates power. That power is only welfare-enhancing when the underlying measurements are more accurate than the judgment they replace.

Unintended Consequences: Designing Against Gaming While Preserving Incentives

The literature on performance measurement is a catalog of pathologies: teaching to the test, threshold gaming, output substitution, cream-skimming, ratchet effects, and outright data manipulation. These are not implementation failures but predictable equilibrium responses to high-powered incentives operating on incomplete measures. Sophisticated performance budgeting design begins by taking these responses seriously and structuring institutions to blunt them.

Ratchet effects illustrate the dynamic complexity. When current performance sets baselines for future targets, rational managers strategically suppress performance to preserve future slack. Soviet-era planning famously suffered from this dynamic, but it appears in contemporary systems whenever year-over-year improvement targets are naively applied. Solutions include yardstick competition against peer organizations, absolute rather than relative benchmarks, and multi-year contracts that reduce the frequency of target renegotiation.

Threshold effects create particularly sharp distortions. When funding depends on crossing specific performance thresholds, agencies concentrate effort on cases near the threshold while neglecting cases far above or below. Hospital quality measures that trigger penalties at specific mortality rates, or school accountability systems keyed to proficiency cutoffs, have generated substantial empirical evidence of this pattern. Continuous rather than discontinuous incentive schedules mitigate but do not eliminate the problem.

Multi-tasking distortions require structural rather than parametric responses. When some dimensions of performance are measured and others are not, uniform low-powered incentives across all dimensions often outperform high-powered incentives on measured dimensions alone. This is a counterintuitive result from mechanism design theory but has important implications: sometimes the optimal performance budget is a less aggressive performance budget, precisely because unmeasured dimensions matter.

Finally, institutional design can leverage information from multiple sources to constrain gaming. Combining administrative data with beneficiary surveys, peer review, and randomized audits creates cross-checks that raise the cost of manipulation. The goal is not to eliminate strategic behavior—which is impossible—but to raise its cost enough that honest performance improvement becomes the dominant strategy.

Takeaway

Gaming is not a bug in performance systems; it is the predictable response of rational agents to incomplete contracts. Good design does not eliminate it but redirects strategic effort toward genuine improvement.

Performance budgeting is neither the accountability revolution its proponents promised nor the bureaucratic failure its critics diagnose. It is a set of institutional tools whose welfare implications depend entirely on design choices about information quality, coupling intensity, and incentive structure. Treating it as a monolithic reform obscures the specific configurations that determine whether it enhances or degrades public sector performance.

The theoretical framework developed here suggests concrete design principles: parsimonious indicator systems with external validation, coupling intensity calibrated to measurement quality, and institutional structures that anticipate and channel strategic behavior. These principles do not guarantee success, but they identify the necessary conditions under which performance budgeting can plausibly improve welfare.

The deeper insight is that performance budgeting is fundamentally an exercise in mechanism design under incomplete information. Progress requires taking the theoretical structure of that problem seriously, rather than pursuing measurement for its own sake or abandoning the enterprise when initial implementations disappoint.