Skip to content
Preprint

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

Aug 2026 · 1 citation · 27 references
Computer Science

TL;DR

RLCascadeRouter is a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop''and model selection, and uses trajectory returns and advantages to directly optimize the performance-cost objective.

Abstract

The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradigms are inflexible: one-shot routers commit before observing responses, whereas conventional cascades stop adaptively but follow a fixed model order. Cascade routing removes both restrictions by reconsidering whether to stop or invoke another model after each response. Current methods use a predict-then-optimize pipeline estimating response quality and future model utility. However, prediction loss for quality or utility is not equivalent to routing-decision loss. A lower prediction error does not necessarily yield a better action; a small boundary-crossing error can reverse a ``stop''or model-selection decision. Therefore, we propose RLCascadeRouter, a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop''and model selection. It uses trajectory returns and advantages to directly optimize the performance-cost objective. Its Cascade Policy Network models candidate complementarity for model selection and remaining-action value for stopping, eliminating independent post-hoc response-quality estimators. Evaluated across ten LLMRouterBench benchmarks with thirteen LLMs, RLCascadeRouter outperforms strong baselines and achieves superior performance-cost trade-offs. It incorporates unseen models without retraining, and ablation studies validate both policy components.

View source

Similar papers

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

TRACE-Router is presented, a task-level routing framework that aligns routing with the unit of supervision, and learns routing policies that adapt to the workload while avoiding explicit task-complexity estimation.

Ritik Raj, Souvik Kundu, Sarbartha Banerjee et al. · 1 citation
Jul 2026

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

Quadrant-weighted Sampling for Length-aware Policy Optimization (QLPO), a simple resampling-based variant of GRPO that introduces implicit length control without modifying the reward function, suggests that structured resampling provides an effective and robust approach to efficient reasoning.

Si-Wei Chen, Si-Qi Chen, Xupeng Miao et al. · 0 citations
Preprint Jul 2026

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B, achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in a comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.

Cheng-Bo Liu, Li-Fang Zhou, Ruijie Yan et al. · 0 citations
#natural language process... Preprint Sep 2026

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the fina...

Shi-Qi Yan, Chao-Hong Tan, Qian Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.