Skip to content
Open access

Embodied Scene Rearrangement Planning

Aug 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 12839-12846 · 0 citations · 31 references
Computer Science

TL;DR

Experimental results demonstrate that current methods struggle to complete the ESRP task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning.

Abstract

This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top-down target layout. Unlike prior rearrangement tasks, Embodied Scene Rearrangement Planning (ESRP) precludes global state access and introduces mutual object occlusions, reflecting the practical constraints of real-world robotic deployment. These factors make aligning partial egocentric observations with the global target layout particularly challenging for long-horizon planning. To facilitate research, we present ESRP-Bench, a comprehensive benchmark built on OmniGibson featuring over 5,400 scene pairs and 8,200 objects. We define three multi-level metrics to evaluate rearrangement quality and provide four baselines: a hierarchical task-and-motion planning method, a vision-language-model-based method, and two learning-based approaches (IL and RL). Experimental results demonstrate that current methods struggle to complete the task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning. This work serves as a stepping stone toward deploying intelligent agents in real-world scenarios.

Read PDF

Similar papers

Preprint Sep 2026

Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge

GRAB-TAMP is introduced, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency.

N. Vijayakumar, Nav Singhal, G. Varma et al. · 0 citations
Preprint Sep 2026

An Information-Space Perspective to Scene Graph Sufficiency for Robotic Task Planning

Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general,...

Basak Sakçak, Francesco Verdoja · 0 citations
Open access Sep 2026

Adaptive Task Planning for Long-Horizon Robotic Manipulation Based on Video Priors and Dynamic Scene Graphs

Robots are now expected to execute increasingly complex long-horizon tasks in unstructured environments. Despite the strong potential of pretrained Vision-Language Models (VLMs) in task planning, their direct application to robotic manipulation is hindered by logical reasoning deviations and inadequate geometric scene...

Guang-Hui Ma, Jia-Hui Guo, Xin-Hua Tang et al. · 0 citations
Preprint Aug 2026

Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance

Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, where a key challenge is selecting target destinations that align with human intentions or preferences. We focus on this challenge in the context of Virtual Placement (VP), the task of identifying all plausib...

Amir Belder, Goncalo Dias Pais, R. Vivanti et al. · 1 citation
Preprint Aug 2026

EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static...

Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction

RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, is introduced, and Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success are introduced.

Tian-Yi Wang, Jia-Zhou Chen, Yiming Xu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.