Skip to content
Preprint

SABRE: A Multi-Agent Approach for Selecting Out-of-Distribution Detectors Under a Budget

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work introduces SABRE (Selective Agentic Budgeted Reliability Ensemble), which replaces this fixed choice with per-regime selection at inference in post-hoc out-of-distribution detection for vision-language models, and shows reliability must be established at deployment rather than assumed from a benchmark.

Abstract

Post-hoc out-of-distribution (OOD) detection for vision-language models assumes that a detector chosen on a benchmark stays reliable once deployed. We show this fails across domains: on a single frozen encoder, a detector that leads in one domain can invert in another, scoring in-distribution inputs as more anomalous than genuine outliers, and the best detector changes from domain to domain, so no fixed choice is reliable throughout. We introduce SABRE (Selective Agentic Budgeted Reliability Ensemble,) which replaces this fixed choice with per-regime selection at inference. Three language-model agents reason over a library of post-hoc detectors under a bounded query budget: a Selector chooses which detector to consult next, a Reporter consolidates the evidence for each input, and an Analyst calibrates detector reliability on a small labeled sample held out from the deployment domain and disjoint from the test data, weighting selection and aggregation without ever observing a scored input's label. The library includes four multimodal density detectors we propose. Inferring the operating regime from data, SABRE tracks the strongest detector in each domain without prior knowledge of it, recovering reliable detection where a conventional detector inverts and converging to that detector where it is sound. A component analysis shows the agents are complementary: the Reporter's feedback yields consistent gains, and the Analyst's calibration is decisive against inversion, ruling out unreliable detectors so that aggregation no longer cancels the sound ones. Since no fixed rule can be trusted across domains, reliability must be established at deployment rather than assumed from a benchmark, and SABRE shows this can be done automatically.

View source

Similar papers

Jul 2026

Level, Sharpness, and Corpus: Why Zero-Shot OOD Detector Rankings Do Not Transfer

The Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles is introduced.

I. M. De La Jara, Cristian Rodriguez-Opazo, Stephen Gould et al. · 0 citations
#artificial intelligence Preprint Aug 2026

JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning

This work proposes the first framework to equip a single compact judge with multi-agent panel deliberation capability at single-model inference cost, and introduces AdaReward, an adaptive multi-reward RL algorithm that dynamically rebalances reward component weights as different objectives saturate at different rates d...

Yi-Yue Qian, Shi-Nan Zhang, Huan Song et al. · 0 citations
Jul 2026

Best-of-Evidence: Best-of-N Selection under Partial Verification

Best-of-Evidence (BoE) is introduced, an inference-time selection framework that keeps the BoN candidate pool fixed, represents reusable claims with a signed candidate--factor graph, and allocates a limited budget to evidence actions that can change the final choice.

Ce Zhang, Teng Fang, Yuxia Wang et al. · 0 citations
Preprint Sep 2026

When is Test-Time Adaptation Identifiable From Unlabeled Evidence?

The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.

Kartik Jhawar, Li-Po Wang · 1 citation
Jul 2026

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.

Qiushi Sun, Kanzhi Cheng, Yian Wang et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.