Every major military conducts elaborate exercises to test readiness, validate doctrine, and demonstrate capability. Billions are spent annually on wargames, field problems, and multinational maneuvers. Yet the historical record reveals an uncomfortable pattern: units that excel in exercises frequently underperform in actual combat, while organizations that struggle in peacetime training sometimes prove remarkably effective under fire.
This disconnect is not incidental. It reflects structural features of how modern militaries design, execute, and evaluate their training. Exercises are shaped by budgetary constraints, safety requirements, political sensitivities, and career incentives that systematically distort the conditions being simulated. The result is a training system that measures something real, but not necessarily what commanders think they are measuring.
Understanding these limitations matters beyond military circles. Exercise assessments inform procurement decisions, alliance commitments, and strategic planning at the highest levels. When policymakers treat exercise results as reliable predictors of combat outcomes, they build strategy on a foundation the training system was never designed to provide.
The Realism Ceiling
Combat is defined by conditions exercises cannot ethically or practically reproduce. Live ammunition creates lethal consequences, so units use blanks, laser systems, or umpire adjudication. Casualties are simulated through cards or notional assessments, meaning no soldier truly grapples with the psychological weight of watching comrades die. Fatigue is bounded by safety protocols. Weather delays halt operations that in war would proceed regardless.
Beyond physical safety, political constraints impose their own ceiling. Exercises involving allies must accommodate national caveats and diplomatic sensitivities. Training on domestic terrain limits maneuver space, aviation profiles, and munitions employment. Environmental regulations restrict where armor can drive and where artillery can fire. Each constraint is individually reasonable; collectively they construct a battlefield that does not exist.
Cost imposes the final compression. Sustained high-intensity operations for weeks are prohibitively expensive, so exercises compress timelines, pre-position logistics, and pre-script critical vignettes. Units rarely experience the cumulative degradation of extended operations, the collapse of supply chains under attack, or the confusion of communications systems overwhelmed by electronic warfare at realistic scale.
The result is a training environment optimized for what can be safely and affordably simulated, not for what combat actually demands. Van Creveld's observation about logistics applies equally here: the friction that defines war is precisely what peacetime organizations are structured to minimize.
TakeawayEvery safety and cost constraint that makes an exercise possible also makes it less predictive. The gap between training and combat is not a bug to be eliminated but a permanent feature to be understood.
The Incentive Distortion
Exercises are not neutral evaluations. They are events where careers are made and broken, where units compete for reputation, and where commanders demonstrate the effectiveness of doctrines they have institutional reasons to defend. These incentives shape exercise outcomes independent of underlying capability.
Consider the classic case of the opposing force. When exercise designers create scenarios where the enemy is defeated on schedule, they may be validating friendly doctrine or they may be constraining the adversary to ensure the desired training objectives are achieved. Aggressor units that consistently defeat rotating brigades have historically been reined in when their success embarrassed institutional narratives. The signal being measured becomes contaminated by the signal being desired.
Officers being evaluated understand these dynamics intuitively. Risk-taking that might prove decisive in war can end a career in peacetime if it produces an equipment mishap or an unfavorable assessment. The rational response is to execute doctrine cleanly, avoid initiative that exceeds authorization, and produce outcomes evaluators expect. Combat rewards different behaviors entirely.
Institutional politics compound the problem. Services defending budget shares design exercises that validate their platforms. Coalition partners avoid scenarios that expose interoperability failures. Contractors supporting training systems have interests in demonstrating those systems work. The exercise thus becomes a negotiated production rather than an experimental test.
TakeawayWhen the observer, the observed, and the outcome all serve institutional interests, the exercise measures institutional health more than combat capability. Ask who benefits from the result before trusting it.
Reading Exercises Correctly
Recognizing these limitations does not make exercises worthless. It changes what they should be used for. Exercises remain excellent for developing procedural competence, testing command and control architectures, building unit cohesion, and identifying logistical bottlenecks. They are poor at predicting who will win a war.
The most useful exercise assessments focus on process rather than outcome. Did staff functions integrate effectively? Did communications architectures survive contested conditions? Did logistics keep pace with maneuver? These questions produce learnable answers. The question of whether Blue Force defeated Red Force in a particular scenario tells you almost nothing generalizable.
Sophisticated militaries increasingly separate certification exercises, designed to validate readiness against defined standards, from experimental exercises, designed to stress systems until they break. The Israeli and American forces that have integrated free-play adversaries with genuine authority to defeat rotating units generate more diagnostic information than scripted validations. The cost is institutional discomfort when favored units lose.
For policy analysts, the implication is caution. Exercise performance should inform hypotheses about capability, not conclusions. Combat effectiveness depends on variables including adaptation under fire, small-unit initiative, and organizational resilience that peacetime training systematically obscures. The organizations that fight best are often those that treat exercises as questions rather than answers.
TakeawayExercises are diagnostic instruments, not predictions. Use them to find what breaks in your system, not to confirm what works. The unit that loses a wargame gracefully may be learning more than the one that wins it cleanly.
The systematic gap between exercise performance and combat outcomes is a permanent feature of military organizations, not a solvable problem. Safety, cost, and political constraints will always shape what can be trained, and institutional incentives will always shape how training is evaluated.
Militaries that understand this design their training accordingly, using exercises to develop capability and diagnose weakness rather than to predict victory. Those that mistake exercise scores for combat readiness build strategy on measurements the system was never engineered to provide.
This tension between what can be trained and what must be fought is one of the enduring problems of military organization. It rewards intellectual humility from commanders, skepticism from analysts, and honesty from institutions willing to learn from what they cannot fully simulate.