A reinforcement-learning model with a continuous action space that integrates sound synthesis during training and live performance is introduced and five complementary reward functions for matching target sounds are proposed.
The proposed RL-based dynamic control system successfully transforms score elements into performance actions which enable robots to deliver expressive music performances.
This course presents multimodal control as the means to unlock that latent capability of motion, identity, environmental sound, lip dynamics, light, and material that a production workflow requires.
Naomi ken korem, Matan Ben Yosef· Proceedings of the Special I...· 0 citations
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing, and VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects...
Wenxiang Guo, Changhao Pan, Ziyue Jiang et al.· 0 citations
This work introduces the Multi-control Mixed Audio Generation benchmark, a systematic evaluation protocol that measures acoustic fidelity, speech quality, semantic alignment, and temporal accuracy, and establishes MMAG as a comprehensive benchmark for future research.
Zihao Zheng, Xuenan Xu, Jiahao Mei et al.· 0 citations
StrixAE, an agent based on a multimodal large language model (MLLM), outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.
Cheng-Lin Wu, Jun-Jie Wu, Jin-Hang Chen et al.· 0 citations
Synthesizer inversion is challenging for two main reasons: 1) Distinct parameter configurations can produce perceptually similar sounds. 2) Parameter-space losses often fail to reflect rendered audio similarity, while the synthesizer being a non-differentiable black box prevents simple audio-domain supervision. To addr...
Tristan Wu, Daniel Chin, Ju-Nan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.