Skip to content
Open access

Feedback Control for Adaptive Language Model Routing Under Non-Stationary Workloads

Jul 2026 · Journal of Computer Science and Technology Studies · 0 citations · 9 references

TL;DR

The result is an interpretable, stability-analyzed, auto-tuned routing controller competitive with or superior to non- stationary bandits at lower cost.

Abstract

Serving a stream of requests across large language models (LLMs) of differing cost and quality is an online allocation problem, usually framed as multi-armed bandits. We frame it as feedback control: a direct-acting Proportional–Integral–Derivative (PID) controller whose setpoint is the running fleet- average performance and whose bounded output adjusts each model’s allocation share, with requests routed by weighted sampling over the allocation vector. This work contributes (i) a stability analysis ofthe closed loop — bounded-input bounded-output behaviour by anti-windup, exponential convergence of the performance estimates via a Lyapunov function, and a persistent-excitation condition guaranteeing recoverability after a regime change; (ii) a closed-form, analysis-grounded automatic tuning rule requiring no per-dataset search; and (iii) an honest head-to-head against static, round-robin, random, epsilon-greedy, UCB1, Thompson sampling, and the non-stationary bandits Sliding-Window UCB and Discounted UCB, on GSM8K with a checkable exact-match reward, reporting inference cost and request latency alongside quality. Under transient drift the auto-tuned controller is statistically tied with the best non-stationary bandit at lower cost; under a persistent regime shift it significantly outperforms both (p < 0.03). We further show the integral term helps only under a persistent shift — a proportional controller suffices for transient drift — and evaluate robustness to noisy rewards. The result is an interpretable, stability-analyzed, auto-tuned routing controller competitive with or superior to non- stationary bandits at lower cost.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

This work proposes Drift-Aware Sparse Routing (DRS), a nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models.

Cheung-Hao Lee, Patrick Wong · 0 citations
Preprint Aug 2026

Refined Thompson Learning for Adaptive Bandits: Power-Efficient Flexibility Scheduling Across Data Centers

A contextual restless multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions is proposed, demonstrating the economic potential of data-center flexibility as a grid service and highlighting the importance of high-quality, open-source AI...

Zi-Xi Chen, Yifu Ding, Rui-Cheng Ao et al. · 0 citations
Preprint Sep 2026

Message-Level Scheduling for RLNC-Coded Multi-Source Traffic

This paper studies weighted decoding-delay minimization for multiple RLNC-coded message streams that compete for finite processing capacity at a destination. Packet arrivals are exogenous, while the scheduler only determines the processing order of packets already available at the destination. A trace-conditioned offli...

Zhao-Hong Lu, Qing-Yu Liu, Hai-Bo Zeng · 0 citations
Open access Aug 2026

SEAL-MAC: Symmetry-Equivariant Lyapunov Actor–Critic for Queue-Stable MEC Offloading

Mobile edge computing (MEC) must serve rapidly growing populations of latency-critical and energy-constrained devices, yet distributed offloading faces two coupled problems: learned multi-agent policies depend on the arbitrary numerical ordering of edge servers, which wastes training samples and treats physically equiv...

Mingchuan Wu, Jian Lu, Yulin Li · 0 citations
Open access Jul 2026

Distributionally Robust Multi-Timescale Elastic Compute Scheduling for Tail-Latency Controlled Microservices

Elastic compute platforms must provision enough replicas to absorb bursty arrivals while avoiding persistent over-reservation. This paper develops DR-MPC-Elastic, a distributionally robust multi-timescale controller for microservice autoscaling. The method replaces a fixed safety margin with a data-dependent ambiguity...

Yizhou Chen · 0 citations
Conference Jul 2026

FlowGuard: Slack-Aware Overload Control for Multi-Agent LLM Serving

Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...

Ali Zafar Sadiq, Hai-Ying Shen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.