Skip to content

Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

Jul 2026 · arXiv.org · Vol abs/2607.25018 · 0 citations · 67 references
Computer Science

TL;DR

Conformal Cascade (CC), a multi-tier inference framework that uses conformal prediction set size as the deferral rule: accept when the calibrated set collapses to a single answer, defer otherwise, delivers a distribution-free, finite-sample accuracy guarantee.

Abstract

Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Production cascades govern this deferral through a confidence threshold, but LLM confidence scores are miscalibrated, the threshold must be tuned per model pair and per domain, and no setting yields a formal bound on cascade accuracy. We introduce \textbf{Conformal Cascade} (CC), a multi-tier inference framework that uses conformal prediction set size as the deferral rule: accept when the calibrated set collapses to a single answer, defer otherwise. The procedure delivers a distribution-free, finite-sample accuracy guarantee. By a per-tier union bound, the prediction set at the accepting tier covers the correct answer with probability at least $1 - K\alpha$ for any user-specified $\alpha$; under a selection-preservation condition (consistent with, but not strictly implied by, our marginal coverage results), the bound tightens to $1 - \alpha$. We further characterise expected cascade cost as an explicit function of $\alpha$ and the calibration-set acceptance rate. Across 18 multiple-choice benchmarks spanning science, medicine, commonsense, and standardized exams, evaluated on two-tier cascades drawn from four open-weight model families, CC strictly improves over the strongest calibration-tuned heuristic cascade on the majority of family--benchmark pairs, with the largest gains on reasoning-heavy benchmarks where majority vote is unreliable; on easier benchmarks the cascade commits the vast majority of queries to the small model at no accuracy cost. Extension to open-ended generation requires an answer-clustering step that we leave for future work. The method requires no model training and only black-box API access.

View source

Similar papers

#machine learning Preprint Aug 2026

Conformalized Large Language Models under Configuration Shift

It is found that configuration shift consistently erodes CP validity, often driving empirical coverage below the target, and coverage lower bounds are derived that attribute this loss to a discrepancy between calibration and test score distributions.

Yuqicheng Zhu, Jia-Lin Yu, Lin Li et al. · 0 citations
2026

A KL Certificate for Best-of-$N$ Reranking in Language-Model Inference

Best-of-$N$ reranking draws independent candidates from a reference policy and selects the response maximal under a fixed, sample-independent strict total order on outcomes. The selected law may differ substantially from the reference in Kullback–Leibler divergence. Prior work introduced a bounded statistic depending o...

Yu-Tong Zhang, Yao-Ran Yang · 0 citations
Preprint Aug 2026

Online Conformal Prediction Beyond Feedback

This work develops OCP with queries (OCPQ) by adapting the label efficient forecaster of Cesa-Bianchi, Lugosi, and Stoltz (2004) to the authors' setting, and develops OCP with queries (OCPQ) with queries in a way that encourages the learner to output small prediction sets while ensuring that the correct label is covere...

J. Skalse, Edoardo Pona, Osvaldo Simeone et al. · 0 citations
Jul 2026

Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models

Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control (CRC) gives rigorous marginal risk guarantees for selective prediction with abstention, but marginal guarantees do not imply per-group ones: a model can meet the population budget w...

Murilo Salem, Luísa Böhm, Daniel Pontes et al. · 0 citations
#natural language process... Preprint Sep 2026

Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

Large language model (LLM) cascades answer easy requests with a small model and escalate selected requests to a larger model. Most routers prioritize examples on which the small model appears uncertain or likely to be wrong. This proxy ignores a decisive fact: escalation is useful only when the large model corrects the...

Zheyuan Wang, Siyu Li, Peiqiao Song et al. · 0 citations
Jul 2026

Simultaneous Coverage and Efficiency Guarantee in Online Conformal Prediction

Adaptive conformal inference (ACI) of Gibbs and Cand{\`e}s and its variants are the standard approach to online conformal prediction under distribution shift, but they suffer from three fundamental limitations. First, their guarantees control only the \emph{signed} long-run coverage error: persistent miscoverage in one...

Rahul Vaze · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.