Jul 2026· Journal of Computer Science and Technology Studies· 0 citations· 9 references
TL;DR
The result is an interpretable, stability-analyzed, auto-tuned routing controller competitive with or superior to non- stationary bandits at lower cost.
Abstract
Serving a stream of requests across large language models (LLMs) of differing cost and quality is an online allocation problem, usually framed as multi-armed bandits. We frame it as feedback control: a direct-acting Proportional–Integral–Derivative (PID) controller whose setpoint is the running fleet- average performance and whose bounded output adjusts each model’s allocation share, with requests routed by weighted sampling over the allocation vector. This work contributes (i) a stability analysis ofthe closed loop — bounded-input bounded-output behaviour by anti-windup, exponential convergence of the performance estimates via a Lyapunov function, and a persistent-excitation condition guaranteeing recoverability after a regime change; (ii) a closed-form, analysis-grounded automatic tuning rule requiring no per-dataset search; and (iii) an honest head-to-head against static, round-robin, random, epsilon-greedy, UCB1, Thompson sampling, and the non-stationary bandits Sliding-Window UCB and Discounted UCB, on GSM8K with a checkable exact-match reward, reporting inference cost and request latency alongside quality. Under transient drift the auto-tuned controller is statistically tied with the best non-stationary bandit at lower cost; under a persistent regime shift it significantly outperforms both (p < 0.03). We further show the integral term helps only under a persistent shift — a proportional controller suffices for transient drift — and evaluate robustness to noisy rewards. The result is an interpretable, stability-analyzed, auto-tuned routing controller competitive with or superior to non- stationary bandits at lower cost.
This work proposes Drift-Aware Sparse Routing (DRS), a nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models.
A contextual restless multi-armed bandit (CRMAB) framework in which a grid operator requests load reductions without observing internal job-scheduling decisions is proposed, demonstrating the economic potential of data-center flexibility as a grid service and highlighting the importance of high-quality, open-source AI...
Zi-Xi Chen, Yifu Ding, Rui-Cheng Ao et al.· 0 citations
This paper studies weighted decoding-delay minimization for multiple RLNC-coded message streams that compete for finite processing capacity at a destination. Packet arrivals are exogenous, while the scheduler only determines the processing order of packets already available at the destination. A trace-conditioned offli...
Mobile edge computing (MEC) must serve rapidly growing populations of latency-critical and energy-constrained devices, yet distributed offloading faces two coupled problems: learned multi-agent policies depend on the arbitrary numerical ordering of edge servers, which wastes training samples and treats physically equiv...
Elastic compute platforms must provision enough replicas to absorb bursty arrivals while avoiding persistent over-reservation. This paper develops DR-MPC-Elastic, a distributionally robust multi-timescale controller for microservice autoscaling. The method replaces a fixed safety margin with a data-dependent ambiguity...
Yizhou Chen· Journal of Computational Met...· 0 citations
Multi-agent applications increasingly rely on shared large language model backends in the public cloud, where bursty workloads cause requests from different agents to contend for the same LLM instances, leading to long queues, memory imbalance, and severe tail-latency inflation. Existing approaches typically prioritize...
Ali Zafar Sadiq, Hai-Ying Shen· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.