This work introduces SABRE (Selective Agentic Budgeted Reliability Ensemble), which replaces this fixed choice with per-regime selection at inference in post-hoc out-of-distribution detection for vision-language models, and shows reliability must be established at deployment rather than assumed from a benchmark.
Abstract
Post-hoc out-of-distribution (OOD) detection for vision-language models assumes that a detector chosen on a benchmark stays reliable once deployed. We show this fails across domains: on a single frozen encoder, a detector that leads in one domain can invert in another, scoring in-distribution inputs as more anomalous than genuine outliers, and the best detector changes from domain to domain, so no fixed choice is reliable throughout. We introduce SABRE (Selective Agentic Budgeted Reliability Ensemble,) which replaces this fixed choice with per-regime selection at inference. Three language-model agents reason over a library of post-hoc detectors under a bounded query budget: a Selector chooses which detector to consult next, a Reporter consolidates the evidence for each input, and an Analyst calibrates detector reliability on a small labeled sample held out from the deployment domain and disjoint from the test data, weighting selection and aggregation without ever observing a scored input's label. The library includes four multimodal density detectors we propose. Inferring the operating regime from data, SABRE tracks the strongest detector in each domain without prior knowledge of it, recovering reliable detection where a conventional detector inverts and converging to that detector where it is sound. A component analysis shows the agents are complementary: the Reporter's feedback yields consistent gains, and the Analyst's calibration is decisive against inversion, ruling out unreliable detectors so that aggregation no longer cancels the sound ones. Since no fixed rule can be trusted across domains, reliability must be established at deployment rather than assumed from a benchmark, and SABRE shows this can be done automatically.
The Complementary Evidence Guard (CEG), a detector-agnostic wrapper that preserves complementary evidence through a non-compensatory fusion of the base detector, level, and sharpness using only empirical in-distribution percentiles is introduced.
I. M. De La Jara, Cristian Rodriguez-Opazo, Stephen Gould et al.· arXiv.org· 0 citations
This work proposes the first framework to equip a single compact judge with multi-agent panel deliberation capability at single-model inference cost, and introduces AdaReward, an adaptive multi-reward RL algorithm that dynamically rebalances reward component weights as different objectives saturate at different rates d...
Yi-Yue Qian, Shi-Nan Zhang, Huan Song et al.· 0 citations
Best-of-Evidence (BoE) is introduced, an inference-time selection framework that keeps the BoN candidate pool fixed, represents reusable claims with a signed candidate--factor graph, and allocates a limited budget to evidence actions that can change the final choice.
Ce Zhang, Teng Fang, Yuxia Wang et al.· arXiv.org· 0 citations
The result is a practical way to separate two failure modes that are usually mixed together: a weak selector versus an information channel that cannot support the desired decision in the first place.
OSReward is introduced, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories, and an open corpus of reasoning-annotated trajectory judgments for the CUA community, to close the gap in reliable CUA reward at scale.
Qiushi Sun, Kanzhi Cheng, Yian Wang et al.· arXiv.org· 1 citation· ⚡1
This thesis is a distinction that is easy to miss: detecting that such a signal helps on average is not the same as learning to act on it per instance, and a reward-SNR floor governs when the second is even possible, and a reward-SNR detectability floor is explained.
Ying Yuan· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.