Skip to content
Preprint

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

UniMoCa, a representation-driven framework that unifies motion and camera controls in visual space, achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity.

Abstract

Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.

View source

Similar papers

Preprint Sep 2026

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instructi...

Xin Shen, Chengyou Jia, Ke Xing et al. · 0 citations
#generative ai Preprint Aug 2026

MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow Trajectories

This work introduces MotionPhys, a lightweight and interpretable framework that treats sparse motion trajectories as physical evidence rather than relying on appearance artifacts or generator-specific traces and reveals subtle motion inconsistencies that are difficult to capture with conventional visual cues and transf...

Hao He, Hao Tan, Zichang Tan et al. · 0 citations
Preprint Aug 2026

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

This work introduces CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments, and proposes CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference.

Da-Zhao Du, Shiyan Du, Jian Liu et al. · 0 citations
Jul 2026

UniCam: Taming Unified Diffusion Models in Noise Space for Camera-controllable Video Rendering

The UniCam framework is proposed, a unified framework that introduces a temporally coherent stochastic representation, termed CameraNoise, warped from camera intrinsic and extrinsic parameters, which significantly outperforms prior methods in both fidelity and controllability.

Haoyu Zhao, Zuxuan Wu, Yu-Gang Jiang · 0 citations
#computer vision Preprint Aug 2026

4DStreamCtrl: Interactive Video Generation with Online 4D Control

This work shows that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass, enabling interactive 4D-controllable streaming generation for the fir...

Shiqian Li, Chenguo Lin, Zhi-Guang Liu et al. · 0 citations
Preprint Sep 2026

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

This work introduces MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video and develops an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision.

Zi-Jie Zhu, Wei-Ren Cai, Yi-Zhou Wang et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.