Skip to content
Open access

Adaptive Non-Playable Characters with Reinforcement Learning: Training and Shipping a Learned Game Opponent

Aug 2026 · International Journal of Innovative Science & Technology · 0 citations · 11 references

Abstract

Most Non-playable characters (NPCs) in modern video games still rely on scripted logic, finite state machines, and decision trees. This reliance on hand-authored rules means that their behaviour is easy to predict and repeat, which reduces player immersion and long-term engagement. Although reinforcement learning has been proposed as an alternative production method, prior demonstrations are often limited to experimental training environments and seldom implemented in a complete playable prototype. The present study fills this gap by taking a reinforcement-learning-trained NPC through the entire production pipeline, from initial design to a tested, standalone game capable of running on ordinary consumer hardware. We implemented a 3D foraging environment called the Flower Island in Unity, where we trained a hummingbird agent (with a 10-value observation vector, 5 continuous actions and a primary reward based on successful nectar collection) using Proximal Policy Optimization with the Unity ML-Agents toolkit and domain randomization applied throughout training. Training was run until the episodic cumulative reward, monitored in Tensor Board, showed progression from exploratory behaviour to more stable behavioural patterns, followed by a reward plateau, at which point the policy was exported for deployment. Once converged, the policy was exported in ONNX format and executed inside the finished game via Unity's Inference Engine on the Burst CPU backend; in playtesting the exported policy ran in real time without observed frame-rate degradation during playtesting and without inference-related runtime exceptions. This was part of a full title with two different game modes, a full user interface and layered audio system. The resultant policy showed purposeful and diverse foraging behaviour and was able to play against a human player in real time. The game passed all ten functional test cases, and the final acceptance test suite ran without any runtime exceptions. Among the engineering factors considered, integration effort, choice of inference backend, and simplicity of the reward structure exerted the greatest influence on the outcome. Formal quantitative benchmarking against a scripted baseline (reward curves, latency figures, and significance testing) was not captured for this iteration of the project and is identified as the primary direction for follow-up work. These results show that adaptive, learned NPCs are possible on commodity hardware and the pipeline described here provides a reproducible template for future games, simulations and educational tools.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.