Jul 2026· Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Courses· pp. 1-3· 0 citations· 6 references
TL;DR
This course presents multimodal control as the means to unlock that latent capability of motion, identity, environmental sound, lip dynamics, light, and material that a production workflow requires.
Abstract
Audio-video foundation models trained at scale implicitly encode a vast repertoire of perceptual and physical knowledge: motion, identity, environmental sound, lip dynamics, light, and material. The practical bottleneck is no longer what such a model can synthesize, but what a user can ask of it. This course presents multimodal control as the means to unlock that latent capability on demand. Once a creative idea is framed as a task, a pairing of audio and visual conditioning signals with the desired output, a short training run on a single machine teaches the model to expose the corresponding behavior. Because the pretraining covers so much ground, the space of attainable controls is in practice open-ended: almost any capability the model already “knows” can be made addressable. We build the course on LTX-2 [HaCohen et al. 2026], an asymmetric dual-stream audio-video foundation model. Heterogeneous conditioning signals are injected into the appropriate stream and learned with the AVControl framework [Ben-Yosef et al. 2026]. We illustrate the framework with several recent applications, all obtained from the same training recipe rather than from bespoke, task-specific architectures: cross-lingual video dubbing, in which the model jointly generates translated speech and the corresponding lip motion while preserving speaker identity, non-speech audio, and visual context [Chen et al. 2026]; SDR-to-HDR video generation via latent alignment with a logarithmic perceptual encoding [Ken Korem et al. 2026]; and a range of finer-grained controls, active-speaker selection (who is speaking when several people share a frame), background replacement, depth- and pose-conditioned generation, and audio transformations, drawn from the AVControl paper itself [Ben-Yosef et al. 2026]. The result is a blueprint for moving from prompting stochastic generators to teaching pretrained models the specific capabilities that a production workflow requires.
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate...
Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing, and VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects...
Wenxiang Guo, Changhao Pan, Ziyue Jiang et al.· 0 citations
VIBE is introduced, a novel text-and-video-to-music (T+V2M) generation model that leverages a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints and soft, subjec...
Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-1...
Jia-Cheng Hua, Xiao-Kun Feng, Jia-Qi Hua et al.· 0 citations
OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generato...
Jia-Shu Zhu, Yan-Hao Zheng, Rui Tian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.