Skip to content

Data Pyramid for Embodied Manipulation

Jul 2026 · arXiv.org · Vol abs/2607.24744 · 4 citations
Computer Science

TL;DR

This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity.

Abstract

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a"pyramid"spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data. We organize the pyramid around the tension between scalability and robot alignment, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity. We then analyze recent embodied foundation models through the lens of their data recipes, examining how different sources are selected, aligned, and mixed during pretraining. For embodied brain models, vision-language-action models, and world-action models alike, we relate data composition to capabilities in perception, reasoning, planning, action generation, and world prediction. We close by discussing six open challenges: building large-scale tactile datasets, collecting failure and recovery data, developing scalable data-collection pipelines, aligning actions across embodiments, leveraging egocentric data for dexterous manipulation, and designing principled data recipes for robot learning. We hope this work paves the foundation for the design of next-generation embodied systems.

View source

Similar papers

Preprint Aug 2026

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $\pi_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.

GigaBrain Team, Angen Ye, Axiang Sun et al. · 1 citation
Preprint Aug 2026

AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization

AnyWorld is proposed, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations and enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain experiences while preserving the underlying dynamics and object interactions.

Cheng Chen, J. Bai, Jiacheng Wei et al. · 2 citations
Preprint Jul 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

Shuailei Ma, Jiaqi Liao, Xinyang Wang et al. · 7 citations · ⚡1
Aug 2026

Visual Embodied Brain-1.5: Enhanced Perception, Spatial Reasoning and Robot Control in Spaces.

Visual Embodied Brain-1.5 (VeBrain-1.5) is presented, a task-level unified framework that connects multimodal perception and spatial reasoning with robot control through a shared MLLM-compatible decision interface and shows strong adaptability, flexibility, and compositional capabilities compared to existing methods.

Ganlin Yang, G. Luo, Ziyang Gong et al. · 0 citations
Review Aug 2026

The Embodiment Gap in Robot Foundation Models

This survey examines what can be reused across robot embodiments and what must still be implemented on a new robot, and examines recent work through three overlapping research directions: sharing semantics and perception, sharing robot data and interfaces, and learning correspondence across embodiments.

Y. Domae, Keisuke Shirai, Hanbit Oh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.