When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
A scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples and develops an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model.