VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing, and VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects is presented.