Skip to content
Preprint

Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning

Aug 2026 · 0 citations · 42 references
Computer Science

TL;DR

A Dynamic Distribution-Aware Uncertainty Quantification framework (DDA-UQ) is proposed that shifts the paradigm from static mapping to a dynamic distribution-aware process and significantly outperforms state-of-the-art methods.

Abstract

Uncertainty Quantification (UQ) aims to measure the reliability of model predictions, serving as a critical safeguard for deploying Vision-Language Models (VLMs) in safety-critical scenarios. Post-hoc approaches are widely adopted due to their lightweight nature, mapping the outputs of VLMs to uncertainty measures through learnable modules or inductive summarization. However, Post-hoc approaches remain inherently confined to fitting the failure patterns of the source domain, ignoring the dynamic nature of test distributions. To address this challenge, we propose a Dynamic Distribution-Aware Uncertainty Quantification framework (DDA-UQ) that shifts the paradigm from static mapping to a dynamic distribution-aware process. During training, we leverage a Gaussian Mixture Model to model the VVLMs'embedding space and extract distributional evidence, thereby dynamically deriving uncertainty estimates. During inference, the design dynamically responds to changes in the data distribution. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model

This work proposes a principled VLM TTA method called \algname, and theoretically reveals that the InfoNCE loss can be neatly reformulated as a Wasserstein OT formulation, thereby unifying the objectives of the inference and adaptation of VLMs to achieve their mutual benefits.

Qi Yu, Zhichen Zeng, Katherine Tieu et al. · 0 citations
Preprint Aug 2026

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

It is shown that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode the authors term prediction-preserving sharpening, and proposed Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a...

Jing-Yan Jiang, Yaru Sun, Xiao Chen et al. · 0 citations
Preprint Aug 2026

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-represented scenarios. This necessitates the need to evaluate the semantic robustness of perception models; conformance of behavior to high-level requirements over real-...

Nusrat Jahan Mozumder, Divya Gopinath, Corina S. Păsăreanu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps

Practical uncertainty quantification (UQ) for large language models must decide, from a single generation, whether a specific answer should be trusted. Existing methods either sample multiple generations, read only output-token probabilities, or reduce the model's internal computation to a single hidden state. We intro...

Jacopo Dardini, Roberta Calegari · 0 citations
Preprint Aug 2026

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

It is suggested that reasoning validity is better read from state-conditioned motion than from either static states or decontextualized trajectories alone, andlations show that motion, region, and direction provide complementary signals.

Hamed Damirchi, I. M. De La Jara, D. Ranasinghe et al. · 0 citations
Preprint Aug 2026

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

PuRF is introduced, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models that consistently outperforms state-of-the-art methods on ViT-B/32 across five datasets.

Yiwen Liang, Hui Chen, Yizhe Xiong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.