Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these ta...
It is shown that uncertainty in audio-language models depends substantially more on available audio evidence than on question text, and top-1 confidence, entropy, and sampling-based measures are established.
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimiz...
This work proposes a semi-supervised generative model that utilizes both labeled and unlabeled samples in a unified framework and maximizes the likelihood of unlabeled samples to learn a latent space shared with the IB on labeled data.
Yiyang Shen, Wei-Ran Wang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.