Skip to content
Preprint

Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

This work asks whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers through guardrailed calls to the HAProxy Data Plane API.

Abstract

Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server returning HTTP 500s until an operator intervenes. We ask whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers through guardrailed calls to the HAProxy Data Plane API. On a reproducible benchmark with a persistent structural fault built into roughly one-third of a heterogeneous fleet, we sweep 15 open-weight models across five families (0.35B to 35B total parameters; dense, mixture-of-experts, and efficient-sparse architectures), reasoning modes, fleet scales of 3 to 9 backends, and two routing algorithms, totaling 240 runs. We find a capability threshold near 3B active parameters. Below it, LLM policies are typically unreliable and sometimes worse than no policy; above it, every model, regardless of architecture, saturates near an 88% reduction in client-perceived 5xx errors over the static baseline. The threshold is approximate: Gemma 4 E2B clears it with 2B active parameters, while the dense 3B Granite 4.0 Micro does not. The availability gain has costs. Draining concentrates load onto surviving servers, inflating tail latency 2.6 to 2.8 times, and enabling reasoning multiplies token spend roughly tenfold, overrunning the control interval and degrading effectiveness. The efficient operating point is a supra-threshold model in its cheapest non-reasoning mode, wrapped inside deterministic guardrails.

View source

Similar papers

Conference Jul 2026

LP-WRR: Towards Adaptive Performance-Aware Load Balancing

Load balancers in practice often rely on fixed heuristics such as weighted round-robin (WRR) or least connection (LC). Although these methods scale well, they do not capture differences in backend service capacity or runtime performance variations. which can increase tail latency and request drop rates in shared cluste...

Hai Pham Thanh, Dang Hoang Nguyen, Anh Nguyen Tuan et al. · 0 citations

FunPilot: Runtime Performance Diagnosis and Remediation for Serverless Applications with LLMs

FunPilot is presented, a system that enables rapid LLM-assisted diagnosis and remediation for serverless applications that uses an event-driven control loop to diagnose active symptoms, derive control knob updates, and validate remediation decisions while coordinating with the underlying autoscaler.

Changyuan Lin, Ya-Wen Wang · 0 citations
Open access Jul 2026

Contamination-Free LLM Routing on LiveBench Reasoning Tasks: Accuracy-Cost-Latency Tradeoff Learning

Fresh, objectively scored benchmark items can support auditable accuracy-cost-latency routing when features encode verifiable computational structure, and show that fresh, objectively scored benchmark items can support auditable accuracy-cost-latency routing when features encode verifiable computational structure.

Grace Xu · 0 citations
Book Open access Aug 2026

Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies

Anytest, an in-situ black-box testing tool that localizes root causes of transport-layer NPAs on commodity RoCEv2 RNICs and Ethernet switches without re-cabling or hardware modification, and implements Anytest's DPDK-based endpoints, which realize protocol correctness while enforcing μs-level packet timing at the hardw...

Zhaochen Zhang, Jia-Qi Gao, Sheng Cheng et al. · 0 citations
Jul 2026

DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs

DeltaServe is presented, a host-agnostic co-serving design that converts this idle inference capacity into LoRA fine-tuning throughput while preserving inference service-level objectives (SLOs).

Jiaxuan Chen, Jianshu She, Ye Yuan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical...

Da Zhao, K. Sankaralingam, Christos Kozyrakis et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.