Jul 2026· Annual International Computer Software and Applications Conference· pp. 1385-1393· 0 citations· 29 references
Computer Science
Abstract
We present lg-sfDA6-X, a language-grounded strategy-following distributed attentional actor architecture after conditional attention, for multi-agent deep reinforcement learning (MADRL). The proposed architecture aims to enable controllable and coordinated agent behaviors in application systems by leveraging a shared saliency representation that integrates environmental conditions with high-level textual instructions provided by external experts and users. To this end, we introduce language-based destination strategies, which allow agents to adapt their behaviors simply by specifying the target regions of the environment using natural language. We evaluate lgsfDA6-X in an object collection game and analyze how agents modify their cooperative and coordinated behaviors in response to diverse, previously unseen, and compositional instructions by users. Experimental results demonstrate that lg-sfDA6-X effectively grounds linguistic instructions into semantically structured latent representations, enabling robust strategy following and coordination. These findings suggest that lg-sfDA6-X provides a promising approach for achieving flexible and interpretable control of coordinated behaviors in MADRL through natural language instructions.
Experiments on the StarCraft Multi-Agent Challenge benchmark demonstrate that CLRE achieves superior win rates and sample efficiency compared with state-of-the-art MARL methods, showing that the proposed approach enables diverse and efficient cooperation among agents in complex environments.
Tingting Wei, Yan Zheng, Zhang-Ling Wang et al.· Autonomous Agents and Multi-...· 0 citations
The proposed method, which extends previous work on controllability in multi-agent deep reinforcement learning, enables uninstructed agents to adaptively complement overlooked tasks and areas and enables uninstructed agents to implicitly complement the overall work based on the actions of other agents.
Y. Takahagi, Gentoku Nakasone, Yoshinari Motokawa et al.· 2025 IEEE/WIC International...· 0 citations
This work proposes AgenticRag-R1, a RL framework that deeply integrates reasoning, retrieval, and memory via a memory stack and fine-grained action space, supported by hierarchical action-aware rewards and an information-aware trajectory rejection strategy to enable effective long-horizon learning.
Xin-Ke Jiang, Yue Fang, Zhi-Bang Yang et al.· 2 citations
Hierarchical Robotic Control (HiRoC) is proposed, a hierarchical post-training framework that decouples high-level task planning from low-level action execution and aligns the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution.
River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.
Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al.· 2 citations
A Self-Evolutional single-agent/multi-agent Reinforcement Learning (SE-RL) framework that utilizes a Large Language Model (LLM) to design various RL algorithm modules, such as agent model design, reward function, profiling, communication, and state imagination, by leveraging the LLM generating module output or code.
Vincent Fu, Xin-Xin Xu, Weichen Xu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.