Reflection-Augmented Strategy Generation via In-Game Triggering in Large Language Models
Abstract
While large language models (LLMs) have made notable progress in strategy generation and task planning, they still struggle to promptly detect tactical imbalances and revise strategies dynamically in complex multi-agent environments such as StarCraft II. To address this, we propose Reflection-Augmented Strategy Generation via In-Game Triggering (RASG-IT), which monitors abrupt drops in unit hit points (HP) to trigger structured reflection. The reflection results are injected into the current decision prompt to guide LLMs in dynamically refining strategies. We evaluate RASG-IT on the StarCraft II Learning Environment for Large Language Models (LLM-PySC2), focusing on the 2s3z and 3s5z scenarios from the Star-Craft Multi-Agent Challenge (SMAC) benchmark. Compared with the LLM-PySC2 baseline method that does not incorporate in-game triggered reflection, RASG-IT raises DeepSeek-V3’s win rate on the 2s3z task from 0% to 35%, and improves its kill-to-death (K/D) ratio from 0.55 to 0.82. On the more challenging 3s5z scenario, it also increases the win rate from 0% to 5%. These results indicate that RASG-IT significantly enhances the effectiveness and interpretability of strategy generation, offering a promising solution for highly dynamic and complex control tasks.