Skip to content

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

It is shown that uncertainty in audio-language models depends substantially more on available audio evidence than on question text, and top-1 confidence, entropy, and sampling-based measures are established.

Abstract

Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.

View source

Similar papers

Preprint Aug 2026

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

This work compares single-pass estimators, which reuse token probabilities from the original generation, with multi-pass estimators, which obtain additional confidence signals through verification or stochastic forward passes, and finds that aggregating token probabilities over the full sequence captures small but cons...

A. Ma, Lorne Schell, Vin Bhaskara et al. · 0 citations
#natural language process... Preprint Sep 2026

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.

Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al. · 0 citations
Preprint Sep 2026

Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating t...

Yu-Pei Li, Qi-Yang Sun, Mohamed Mady et al. · 0 citations
Preprint Aug 2026

Asymptotic Risk Calibration for Selective Question Answering

A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.

Shufan Lin, Sijin Dong · 0 citations
#natural language process... Preprint Sep 2026

Likelihood Ranking doesn't Scale Like Prompting in LLMs

LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned an...

Alessandro Bondielli, Lucia C. Passaro, D. Bacciu et al. · 0 citations
Review Aug 2026

Uncertainty-Aware Decision Making in Multimodal Large Language Models

This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action.

Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.