Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion. CALM-MAS detects early signals of back-end saturation, and leverages the flexibility of LLM applications to regulate load. During spikes of requests, the system downgrades agent topology and reasoning depth; during low-utilization periods, it allocates additional reasoning effort to maximize task accuracy. Compared with a static serving baseline based on vLLM, CALM-MAS reduces shared-backend tail latency by 77% with the accuracy degradation remaining confined to 6.1 pps relative to the native agent configuration.
Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al.· Proceedings of the 17th ACM...· 0 citations
This work introduces CoLEDS, a method for profiling unlabeled client datasets with minimal computational overhead that yields federatively trained models that are better aligned with individual data distributions and enables appropriate model assignment even for clients that do not participate in federated training.
Boris Radovič, Marco Canini, V. Pejović· Data mining and knowledge di...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.