AI-Supported Adjudication in Educational Wargames: Integrating LLMs for Strategic-Level Simulations
DOI:
https://doi.org/10.34190/ecgbl.20.1.5476Keywords:
AI-supported adjudication, educational wargaming, Large Language Models, strategic-level simulation, DIME/PMESII framework, game-based learningAbstract
Traditional matrix wargaming relies on human expertise to adjudicate outcomes, which often limits scalability and introduces facilitator bias. This paper introduces an approach using AI to support adjudication in educational wargames on the strategic level designed for professional military education (PME). We discuss the application of Large Language Models (LLMs) as an "Adjudication Engine" that processes qualitative player arguments by cross-referencing them with domain data to ensure consistent game world updates. Rather than replacing human judgment, the AI serves as a structured analytical layer that synthesises competing player moves into coherent technical reports across the Diplomacy, Information, Military, and Economic (DIME) dimensions. The study is based on the "Arctic Hybrid 2026" Educational Matrix Wargame, conducted at the Bundeswehr Command and Staff College in February 2026, where 18 participants in six teams navigated a fictitious Arctic crisis scenario involving hybrid power projection, based solely on publicly available data. The design integrates a stochastic D20 dice mechanic with qualitative DIME move sets and PMESII (Political, Military, Economic, Social, Infrastructure, Information) effect analysis. Seven AI-steered non-player characters were configured with distinct strategic profiles and behavioural parameters. We detail the "Adjudication Matrix" workflow: how player inputs and dice results are transformed by the AI into a coherent technical report, maintaining the simulation’s "State of the World" across an initial "Move 0" and two additional game rounds. The use of grounded AI systems — most centrally NotebookLM, operating exclusively on curated source corpora continuously updated with game-generated data — proved critical for containing hallucination risk within operationally acceptable bounds. By semi-automating technical adjudication, the game allowed for greater complexity of actors, more rapid turns, and deeper immersion. We observe how structured AI feedback helped participants better understand nonlinear consequences of strategic decisions, including cascading alliance failures and hybrid fait accompli tactics. The paper discusses both potential and limitations of LLM-based adjudication, including the necessity of human oversight and hallucination management in security-sensitive educational contexts.