Skip to content
Preprint

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Aug 2026 · 0 citations · 48 references
Computer Science

TL;DR

Vorch-Omni is presented, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation that supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing.

Abstract

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.

View source

Similar papers

Preprint Aug 2026

EchoWM: Open and Enterable Omnimodal World Models

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes,...

Song-Chun Zhang, Yao-Wei Li, Junhao Zhuang et al. · 5 citations · ⚡1
#artificial intelligence Preprint Sep 2026

OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents

Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce...

Rui-Xun Liu, Yuxuan Wang, Jia-Cheng Xie et al. · 1 citation · ⚡1
Preprint Aug 2026

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Human-centric audio-visual generation spans several closely related tasks: animating a person from driving speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references. Existing systems commonly solve these tasks with separate models, even thou...

Yang Ding, Hao-Ran Yu, Xin Ma et al. · 0 citations
Preprint Sep 2026

ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts

Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 task instances with 2-1...

Jia-Cheng Hua, Xiao-Kun Feng, Jia-Qi Hua et al. · 0 citations
Preprint Aug 2026

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

ST-Omni-R1 is proposed, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning, and results on three public spatial-audio benchmarks indicate that its learned spatial and motion r...

Zhi Zeng, Cheng Zhang, Ze-Sheng Yang et al. · 2 citations
Preprint Aug 2026

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.

Nan Duan, Hao-Yang Huang, Wei-Yang Jin et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.