Skip to content

EgoMotion: Hierarchical Vision-Language Learning and Diffusion for Egocentric Motion Generation.

Sep 2026 · IEEE Transactions on Pattern Analysis and Machine Intelligence · Vol PP, pp. 1-14 · 0 citations
Medicine

TL;DR

EgoMotion is proposed, a two-stage framework for vision-language-guided egocentric motion generation that achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.

Abstract

Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely underexplored due to the inherent complexity of first-person perception. In this work, we investigate Egocentric Vision-Language (Ego-VL) motion generation. This task requires synthesizing 3D human motion conditioned jointly on first-person visual observations and natural language instructions. We identify a critical optimization challenge in effectively transferring vision-language understanding to motion generation. Directly optimizing vision-language semantic learning and kinematic motion synthesis in an end-to-end manner can lead to optimization interference, limiting the quality of generated motions. To address this challenge, we propose EgoMotion, a two-stage framework for vision-language-guided egocentric motion generation. In the first stage, a vision-language model (VLM) learns motion-aware semantic representations from multimodal inputs through autoregressive motion token prediction. In the second stage, these learned VLM representations serve as expressive conditioning signals for a diffusion-based motion generator. By performing iterative denoising within a continuous latent space, the generator synthesizes physically plausible and temporally coherent trajectories. Extensive evaluations demonstrate that EgoMotion achieves state-of-the-art performance and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.

View source

Similar papers

Preprint Aug 2026

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

CL4D is the first foundational 4D vision encoder that directly operates on dynamic point clouds, trained with a contrastive learning objective to align spatio-temporal geometric representations with natural language descriptions, and 4DVLM, a 4D vision-language model that conditions language generation on dynamic geome...

K. Hewagamage, I. Senavirathne, S. Amarasinghe et al. · 0 citations
Preprint Sep 2026

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...

Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al. · 1 citation
#small language model Review Aug 2026

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI

This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill...

M. Zamani, Fatemeh Ziaeetabar · 0 citations
#artificial intelligence Preprint Sep 2026

Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World

Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-mod...

Guo-Cun Wang, Kenkun Liu, Guo-Rui Song et al. · 0 citations
#machine learning Preprint Sep 2026

MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies

Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared g...

Jing-Qi Wang, Yan Wang · 0 citations
Preprint Sep 2026

Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models

This work introduces Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties and proposes the TempoVista framework, featuring the Kinematic-GSPO algorithm.

Jiayu Ding, Zhuo-Dong Liu, Lei Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.