Skip to content
Open access

Coverage-Aware Guidance for Novelty-Driven Exploration in Automated Game Testing Under Sparse-Reward 3-D Environments

2026 · IEEE Access · Vol 14, pp. 137393-137407 · 0 citations · 49 references

Abstract

Automated testing of high-fidelity 3D games requires agents to explore large state spaces under sparse rewards and execute context-dependent interactions that expose faults. Random network distillation (RND) promotes exploration by rewarding observations that remain difficult for a predictor network to estimate, but prediction error does not explicitly represent exercised state-action contexts. We propose RND+CAE, which combines RND with coverage-guided adaptive exploration (CAE). CAE maintains visitation counts over environment-adapted coverage abstractions and provides a one-step recovery bonus when structured coverage stops expanding. We evaluate the framework in a randomized maze for spatial exploration and a partitioned arena for interaction-dependent fault discovery. RND remains a strong maze baseline, while RND+CAE achieves comparable cumulative coverage with lower observed variability. In the Arena, RND+CAE obtains the highest mean final unique bug count and bug-discovery area under the curve (AUC) among the main methods. An ablation shows that spatial-only guidance matches Full RND+CAE in final fault breadth but accumulates faults more slowly, whereas removing stagnation recovery reduces bug AUC and the proportion of runs reaching five bugs. These results indicate that geometric coverage strongly influences eventual fault breadth in the fixed Arena and that Full CAE’s clearest additional contribution is sustained late-stage discovery through stagnation recovery. The independent benefits of interaction-oriented keys are not established consistently. The framework may also support reliability validation in complex disaster-response simulations.

Read PDF

Similar papers

#software testing Preprint Aug 2026

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

This work introduces Risa (Routing-Informed Steering and Arbitration): within trajectories, routing encourages diverse exploration and controlled convergence during patch commitment; across separately sampled trajectories, agreement at informative patch positions selects a final candidate.

Kang Chen, Junjie Nian, Yi-Xin Cao et al. · 0 citations
#machine learning Preprint Sep 2026

Learning to Plan from Random Exploration

Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer...

De-Qian Kong, Guang-Yan Sun, Sheng Cheng et al. · 0 citations
Open access Aug 2026

Graph-Guided Item Randomizers: Formal Models, Search Strategies, and Metrics for Procedural Game Worlds

This paper formulates key-item placement in adventure-style games as a reachability problem on directed graphs with rule-constrained edges, introducing formal algorithms that support progression feasibility while enabling controlled randomness. Three placement procedures are defined in this paper: Random Fill, Forward...

Monisha Rengaraj · 0 citations
#machine learning Preprint Sep 2026

Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards

Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed trajectory can reveal which earlier choice...

Fan-Chao Chen, Heng-Yu Fu, Shivaram Venkataraman et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

Online agent deployments accumulate execution trajectories at massive scale and behavioral diversity, for which predefined annotation criteria hardly exist. Extracting useful evidence therefore demands costly manual annotation or verifier signals that fail to scale, leaving valuable evidence buried among redundant, inc...

Si-Yuan Liu, Fan Yu, Dongyu Ru et al. · 0 citations
#natural language process... Preprint Sep 2026

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

This work proposes ArenaFlow, a hierarchical credit propagation framework for open-ended agent reinforcement learning that leverages tournament-based relative ranking to derive trajectory-level reward signals and propagates trajectory-level advantages to high-confidence pivotal steps according to tournament survival de...

Qiang Zhang, Rui-Xue Ding, Fanrui Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.