Skip to content
Preprint

Seeing the World and the Self from Egocentric Video

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation.

Abstract

Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods lack explicit scene geometry and often depend on external trajectories. Joint recovery is challenging because the two tasks exhibit asymmetric visibility and require different prediction paradigms. The largely visible scene supports deterministic geometric regression, whereas the severely occluded body requires generative motion inference. We therefore propose RESELF (REconstructing the Scene and the sELF), a unified framework that couples deterministic metric geometry reconstruction with geometry-conditioned motion generation. RESELF adapts a geometry foundation model pre-trained on large-scale exocentric data to egocentric video using frame-wise scale and relative-pose consistency objectives. The resulting camera trajectory and latent geometric features condition a diffusion model that recovers the wearer's motion. A subsequent closed-loop kinematic feedback stage further refines the camera head while preserving the reconstructed scene geometry. To support training and evaluation, we curate EE4D-JSM from EgoExo4D by aligning egocentric video, sparse metric scene geometry, camera trajectories, and full-body motion annotations. Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation. Code, models, and datasets will be available at https://ka1guan.github.io/RESELF/.

View source

Similar papers

Preprint Aug 2026

ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

ACE-Ego-Hand is introduced, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder, offering a scalable path from everyday human video to robot manipulation data.

Yu-Fei Liu, Xixi Wang, Hao Li et al. · 2 citations · ⚡1
#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
Preprint Oct 2026

EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to sepa...

Hong-Ming Fu, Jing-Cheng Shi, Wen-Jia Wang et al. · 0 citations
Preprint Oct 2026

4Director: Controlling Video World Models with Rigid 3D Geometry

Precise control over camera and object motion is essential for professional video production. Existing methods control objects only coarsely, through image-plane cues that are ambiguous in depth and rotation or through 3D tracks and blobs that lack complete geometry and lose consistency across viewpoint changes. We int...

Wei Cao, Hao Zhang, Vikram Voleti et al. · 0 citations
Aug 2026

Dynamic View Synthesis from Monocular Videos via Motion-aware Gaussian Splatting.

This paper proposes a semantics-guided scene decoupling module that separates Gaussian primitives into static and dynamic components based on motion vectors, and introduces a motion-aware densification module for motion compensation, which alleviates the incomplete rendering of dynamic objects caused by insufficient sp...

Chulin Zhao, Xue Wang, Guo-Qing Zhou et al. · 0 citations
Preprint Sep 2026

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

This work introduces MINT (Minting IN-the-Wild Trajectories), a foundation model for world-space hand motion reconstruction from ego-centric RGB video and develops an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision.

Zi-Jie Zhu, Wei-Ren Cai, Yi-Zhou Wang et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.