Estimating Cognitive Load via Process-Based Telemetry for Adaptive Serious Games
DOI:
https://doi.org/10.34190/ecgbl.20.2.5489Keywords:
Serious games, Cognitive load, Telemetry, Interaction metrics, NASA-TLX, EducationAbstract
Traditional cognitive load measures interrupt gameplay: the NASA-TLX requires time to complete, physiological sensors demand lab conditions. Neither integrates cleanly into a serious game session. This study tests whether sixteen process-based interaction metrics, captured automatically through in-game telemetry, can track perceived cognitive load without interrupting play. Thirty university students played LUTLAB (a puzzle-based serious game for learning FPGA hardware concepts) across ten levels. The game's telemetry logged mouse movement speed and distance, movement pause patterns (contemplation and hesitation frequencies, pause ratios), duplicate game states, test-button and tool usage, click frequency and burstiness, help-seeking behavior, and dwell time in task rooms. Players completed the Raw Task Load Index (Raw-TLX) after each level. Spearman rank correlations were computed between every metric and Raw-TLX scores across 150 level-observations. One metric held up: dwell time in task rooms correlated with perceived load at (ρ = 0.48–0.66, all p < 0.01) in every one of the five analyzed levels. Active mouse movement speed reached stronger correlations in individual levels (up to ρ = −0.67 in an extended analysis) but did not hold globally. Duplicate game states and contemplation pause frequency predicted load primarily in Level 3, the first level where players could freely interact with game elements. Most metrics showed weak or no correlation with Raw-TLX scores. What this means for adaptive game design is direct: there is no single behavioral signal that works uniformly across all task types. A metric's predictive value is tied to the structure of the specific task, not to any general property of the metric itself. Cognitive load indicators cannot simply be applied across an entire game; they must be selected and validated within the context of each module. This study provides systematic empirical evidence for this context-dependency across multiple levels of the same game, establishing a basis for task-aware behavioral sensing in adaptive serious games.