Skip to content
Preprint

Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

This work proposes ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification), an accurate and efficient framework that detects ambiguity via logical conflict, and presents QuireQA, a 4,703-query benchmark spanning factoid, non-factoid, and ill-formed queries.

Abstract

How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification leads to answering the wrong interpretation or unnecessary clarification. However, existing methods conflate answer diversity with ambiguity, leading to inaccurate predictions. They also process queries uniformly, resulting in wasteful computation. We propose ARCHIVE (Ambiguity Recognition via Cascaded Hypothesis Inspection and Conflict Verification), an accurate and efficient framework that detects ambiguity via logical conflict: a query is ambiguous when its valid answers cannot all be true under a single interpretation. ARCHIVE combines a lightweight early-exit encoder for surface-detectable cases with a conflict reasoning module that models logical relations among answers, reinforced by an invariance objective for robustness to noisy answer sets. We present QuireQA, a 4,703-query benchmark spanning factoid, non-factoid, and ill-formed queries. Experiments show ARCHIVE outperforms competitors, improving F1-amb by up to 10.4% and F1-unamb by up to 21.6%, while operating 16$\times$ faster than the best competitor.

View source

Similar papers

#natural language process... Preprint Sep 2026

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.

Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al. · 0 citations
Preprint Aug 2026

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.

Kuangzhao Yang, Ziliang Zhao, Zhi-Cheng Dou · 0 citations
Book Open access Sep 2026

Efficient Ambiguity Resolution via Information-Guided Clarification for Text-to-SQL

Text-to-SQL technology enables users to query relational databases using natural language, but ambiguity in user questions remains a major obstacle: non-expert users naturally produce ambiguous queries, yet current Text-to-SQL systems often ignore ambiguity or rely on interaction without a principled strategy, resultin...

Lu-Yu Qiu, Jia-Ning Li, Yuan-Feng Song et al. · 0 citations
Preprint Aug 2026

Asymptotic Risk Calibration for Selective Question Answering

A-CRC-QA is a post-hoc calibration framework for uncertainty-aware selective question answering that reformulates selection-conditioned error control as a linear expectation constraint and applies a monotonized empirical-risk calibration procedure inspired by conformal risk control.

Shufan Lin, Sijin Dong · 0 citations
Preprint Aug 2026

When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

CROWN-QA is introduced, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants.

Byoungjae Min, Kennedy Edemacu, Sae-Hong Cho et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Semantic Uncertainty Quantification Needs Factual Equivalence

Semantic uncertainty quantification for large language models rests on a common template: sample several answers, measure how much they agree, and treat disagreement as uncertainty. We first formalize this template as two separate roles: an operator that compares two answers, and an aggregator that combines all pairwis...

Joseph Hoche, Quentin Guimard, Gianni Franchi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.