Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating t...
Yu-Pei Li, Qi-Yang Sun, Mohamed Mady et al.· 0 citations
Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under orig...
Qi-Yang Sun, Xu-Dong Li, Yu-Pei Li et al.· 1 citation
TopoAlign is proposed, a framework that unlocks widely available code repositories as training resources for Math LLMs and decomposes code into docstrings, main functions, and dependency functions, and reassembles these components into analogues that structurally mirror formal statements.
Yu-Pei Li, Philipp Borchert, Gerasimos Lampouras· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.