Abstract Artificial intelligence-based approaches to speech analysis have the potential to assist with objective speech error analysis in aphasia but off-the shelf tools often fail to detect speech errors due to prioritizing ‘fluent transcription’. Speech production errors (dysfluencies) are hallmark diagnostic feature...
J. Vonk, Jia-Chen Lian, G. Kurteff et al.· Brain Communications· 0 citations
A pretrained text-to audiovisual generation model is adapted through source-conditioned feature modulation to jointly learn object addition and removal, and quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition.
Wei-Han Xu, Kan-Jen Cheng, Koichi Saito et al.· 0 citations
HuPER is a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge and is the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions.
Chen-Xu Guo, Jia-Chen Lian, Yi-Si Liu et al.· arXiv.org· 4 citations· ⚡1
TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio, and is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
Akshaj Gupta, Hwi Joo Park, Andrea Guzman et al.· 0 citations
This work systematically decomposes OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluates 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis, showing th...
Akshat Gupta, Jermaine Lei, Alexander Lu et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.