Skip to content
Conference

Congestion-Aware Scheduling for Heterogeneous LLM-Agent Teams

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 2871-2876 · 0 citations · 32 references
Computer Science

Abstract

Coordinating heterogeneous LLM agents under congested online settings is difficult because bursty task arrivals and limited per-agent capacity may induce hotspot overload and severe tail waiting time. This paper proposes an online scheduling method based on candidate-set contraction before assignment. Specifically, tasks are first routed through subscription matching to identify a task-relevant candidate pool, after which layered gating is applied to enforce capability feasibility, historical quality, and real-time load constraints. The remaining candidates are then ranked using a composite score that balances competence and load, with stable tie-breaking introduced to reduce assignment fluctuations under contention. We evaluate the method under a reproducible protocol with both regular and congested regimes. Across benchmarks covering code generation, arithmetic reasoning, and preference-based evaluation, the proposed approach preserves competitive task performance while reducing both mean and 95th-percentile waiting time in congested settings relative to representative linear, flat, and hierarchical baselines. The findings suggest that candidate contraction is a useful strategy for achieving more stable coordination in heterogeneous LLM-agent systems.

View source

Similar papers

Conference Aug 2026

Discrete Modeling and Combinatorial Optimization for Task Scheduling

This paper investigates centralized scheduling of mobile service agents under staggered multi-wave task arrivals, limited service capacity, service-time windows, and cross-wave capacity reservation requirements. A spatiotemporal candidate-arc representation is developed to discretize the continuous scheduling process,...

Miao Shen, Chuan-Fu Guo, Peng Wang et al. · 0 citations
Conference Jul 2026

FlowGuard: Slack-Aware Overload Control for Multi-Agent LLM Serving

Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...

Ali Zafar Sadiq, Hai-Ying Shen · 0 citations
Book Open access Sep 2026

Congestion-Aware Serving of Agentic LLM Applications

CALM-MAS is proposed, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion.

Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al. · 0 citations
Preprint Aug 2026

MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning

This work proposes MARA, which predicts future loss trajectories with conditional flow matching and coordinates compute nodes through a cooperative multi-agent autoregressive policy and reduces remaining-resource prediction error relative to weighted least squares.

Han-Ye Zhao, Mu-Ning Wen, Yong Yu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.