L2AF: Graph-attention modulated learning automata for efficient multi-agent cooperation.
Abstract
Multi-agent reinforcement learning (MARL) is impeded by the combinatorial explosion of joint action spaces and the neglect of complex inter-agent dependencies, which severely hinder effective coordination and accurate credit as signment. To address these challenges, this paper proposes a novel framework termed Lightweight Learning Automata with Feature extraction (L2AF), integrated with a Graph-Attention Modulated Feature Extractor (GAMFE). Specifically, GAMFE synergizes compact MLP embeddings with graph attention networks to explicitly model relational dependencies among agents. This mechanism generates attention-weighted continuous reward signals (S-model) that dynamically modulate the update step size of the automata. Theoretically, via ordinary differential equation (ODE) modeling and Lyapunov stability analysis, we rigorously establish that L2AF achieves local asymptotic stability and converges to the global optimal joint action at an exponential rate. Experimental evaluations on stochastic coordination environments demonstrate that L2AF achieves an AUC of 0.77 ± 0.11, significantly outperforming state-of-the-art baselines such as LA-OCA and EMAQ. Furthermore, ablation studies confirm that the attention mechanism and continuous feedback yield performance gains of 6.9% and 11.6%, respectively, while their integration reduces learning variance by 15%, highlighting the framework's robustness.