A runtime strategy-selection framework in which a large language model (LLM) guides a trained RL policy without modifying its underlying behavior is investigated, demonstrating both the potential and limitations of LLM-guided runtime strategy selection for adaptive multi-agent game AI.
Abstract
Scripted and rule-based non-player characters (NPCs) in combat video games often exhibit predictable behaviors that experienced players can exploit, while reinforcement learning (RL) agents typically retain a fixed policy after training and cannot readily adapt their strategy to different opponents. We investigate a runtime strategy-selection framework in which a large language model (LLM) guides a trained RL policy without modifying its underlying behavior. To demonstrate this, we train five NPC agents with a shared PPO policy in Unity and compare a baseline configuration, in which the policy acts independently, with an LLM-augmented configuration in which a locally hosted Mistral 7B model, accessed through Ollama, reads the live game state every five seconds and assigns one of four tactical tags. We evaluate both configurations against three scripted opponent types across 600 episodes and analyze outcomes using the Mann-Whitney U test. Against a Balanced opponent that changes tactics during an episode, the LLM-augmented agents more than doubled their win rate from 11% to 24% and produced significantly longer episodes. Against an Evasive opponent, the augmented agents achieved a higher win rate and faster kills, although their shorter episode duration did not satisfy the strict hypothesis definition. Against an Aggressive opponent, the LLM's near-constant preference for encirclement was counterproductive. Analysis of 2,430 strategy selections showed that Surround was selected in 83.8% of cases regardless of opponent type, indicating limited zero-shot strategic differentiation at this model scale. These results demonstrate both the potential and limitations of LLM-guided runtime strategy selection for adaptive multi-agent game AI.
Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths, produces policies that generalize more effectively to unseen counterparts.
Senhao Wang, Chenghao Cai, Hai-Tao Hu et al.· 0 citations
Most Non-playable characters (NPCs) in modern video games still rely on scripted logic, finite state machines, and decision trees. This reliance on hand-authored rules means that their behaviour is easy to predict and repeat, which reduces player immersion and long-term engagement. Although reinforcement learning has b...
Imran Maqsood, Namos Khan, Hijab Noor ul Amin et al.· International Journal of Inn...· 0 citations
UnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn trajectories with Python tool calls, and an Evaluation Player that constructs executable verifiers, highlights cooperation among specialized players as a promising path toward self-enh...
Most existing game AI research examines individual mechanisms—Dynamic Difficulty Adjustment (DDA), large language model (LLM)-driven NPC dialogue, and Procedural Content Generation (PCG)—in isolation, leaving open the question of how these subsystems interact when evaluated together against a shared user-experience (UX...
Abhinav, Amandeep, Dharmender Kumar, Suraj, Keshav· International Journal of Adv...· 0 citations
A Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework that utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code.
Vincent Fu, Xin-Xin Xu, Weichen Xu et al.· Proceedings of the 32nd ACM...· 0 citations
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of...
R. Truong, Lance Ying, Samuel Gershman et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Martin Trust Center Managing Director Bill Aulet introduces Dear Dreamer, a free platform for middle and high school students who want to learn about entrepreneurship.
Microsoft Research Blog· microsoft.comSep 30, 2026
Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.