Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-Function
This work proposes MAFIS, a novel method that addresses limitations for both online and offline MAIL settings by building upon the single-agent IQ-Learn framework and introducing the value decomposition network to factorize the imitation objective at agent level, thus enabling scalable training for multi-agent systems.