Competitive Testing of Distributed LLM Systems Based on Multi-Agent Learning with Value Function Decomposition
Abstract
Distributed large language model infrastructures create fundamentally new semantic attack surfaces involving orchestration layers, retrieval pipelines, autonomous agents, moderation systems, distributed memory components, and adaptive guardrails. Existing adversarial testing methods are inadequate for enterprise-scale semantic ecosystems. The combinatorial explosion renders manual prompt engineering unscalable, and stochastic moderation causes static optimization techniques to degrade rapidly. This paper presents a cooperative multi-agent reinforcement learning framework for controlled competitive testing of distributed LLM systems modeled as a decentralized partially observable Markov decision process (Dec-POMDP). Two agents jointly select semantic strategy classes and representation-layer transformations in a coordinated ten by ten action space. The study evaluates Value-Decomposition Networks (VDN) and QMIX against Independent Q-learning and random exploration. The experiment was run for 300,000 epochs per algorithm with probabilistic blocking, replay-buffer training, target networks, and a fixed random seed. The results show that value factorization stabilizes cooperative search under partial observability: QMIX and VDN achieve close asymptotic performance and outperform Independent Q-learning and the random baseline. The framework is positioned as a defensive, auditable testing method with explicit reward formalization, governance constraints, and reproducibility requirements.