Skip to content

Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey

Feb 2026 · arXiv.org · Vol abs/2603.04445 · 21 citations · ⚡ 1 influential · 113 references
Computer Science

TL;DR

This survey provides a systematic analysis of multi-LLM routing and cascading approaches, focusing on systems that route queries across a pool of independently trained LLMs at inference time, and demonstrates that effective multi-LLM routing requires balancing competing objectives.

Abstract

The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent model selection at inference time. While smaller models suffice for routine queries, complex tasks demand more capable models. However, static model deployment does not account for the complexity and domain of incoming queries, leading to suboptimal performance and increased costs. Dynamic routing systems that adaptively select models based on query characteristics have emerged as a solution to this challenge. This survey provides a systematic analysis of multi-LLM routing and cascading approaches, focusing on systems that route queries across a pool of independently trained LLMs at inference time. We cover diverse routing paradigms, including query difficulty, human preferences, clustering, uncertainty quantification, reinforcement learning, multimodality, and cascading. For each paradigm, we analyze representative methods and examine key trade-offs. Beyond taxonomy, we introduce a conceptual framework that characterizes routing systems along three dimensions: when decisions are made, what information is used, and how they are computed. This perspective highlights that practical systems are often compositional, integrating multiple paradigms under operational constraints. Our analysis demonstrates that effective multi-LLM routing requires balancing competing objectives. Choosing the optimal routing strategy depends on deployment and computational constraints. Well-designed routing systems can outperform even the most powerful individual models by strategically leveraging specialized capabilities across models while maximizing efficiency gains. Meanwhile, open challenges remain in developing and evaluating routing mechanisms that generalize across diverse architectures, modalities, and applications.

View source

Similar papers

#machine learning Preprint Sep 2026

Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model...

Muhammad Abdur Rab Siddiqui, Daniel Rojas, Chen Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the overall probability of success. To address...

Wang Wei, Harry Yang, Tiankai Yang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Pretrain Once, Route Anywhere: Towards a Foundation Model for LLM Routing

Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate poo...

Guannan Lai, Han-Jia Ye · 0 citations
Preprint Aug 2026

RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning

RLCascadeRouter is a quality-estimator-free framework that formulates cascade routing as a Markov decision process with actions comprising ``stop''and model selection, and uses trajectory returns and advantages to directly optimize the performance-cost objective.

Shihong Huang, Sheng-Jie Wang, Hong-Yao Ma et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing...

Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis · 0 citations
Preprint Aug 2026

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

This work presents a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing.

Tao Feng, Fangxu Yu, Haozhen Zhang et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.