Preprint
Aug 2026
Training Small LLMs as Spatial Multi-Agent Policies
Training LLM-based multi-agent systems with multi-agent reinforcement learning with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.
Yi Mao, Andrew Perrault
· 0 citations