Skip to content
Book Open access

Congestion-Aware Serving of Agentic LLM Applications

Sep 2026 · Proceedings of the 17th ACM SIGOPS Asia-Pacific Workshop on Systems · 0 citations · 41 references

Abstract

Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an elastic resource, dynamically adjusting the compute profile of admitted tasks to tame congestion. CALM-MAS detects early signals of back-end saturation, and leverages the flexibility of LLM applications to regulate load. During spikes of requests, the system downgrades agent topology and reasoning depth; during low-utilization periods, it allocates additional reasoning effort to maximize task accuracy. Compared with a static serving baseline based on vLLM, CALM-MAS reduces shared-backend tail latency by 77% with the accuracy degradation remaining confined to 6.1 pps relative to the native agent configuration.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.