Aug 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 12839-12846· 0 citations· 31 references
Computer Science
TL;DR
Experimental results demonstrate that current methods struggle to complete the ESRP task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning.
Abstract
This paper introduces Embodied Scene Rearrangement Planning (ESRP), a novel task requiring embodied agents to rearrange furniture in 3D scenes to match a target configuration using only egocentric observations and a top-down target layout. Unlike prior rearrangement tasks, Embodied Scene Rearrangement Planning (ESRP) precludes global state access and introduces mutual object occlusions, reflecting the practical constraints of real-world robotic deployment. These factors make aligning partial egocentric observations with the global target layout particularly challenging for long-horizon planning. To facilitate research, we present ESRP-Bench, a comprehensive benchmark built on OmniGibson featuring over 5,400 scene pairs and 8,200 objects. We define three multi-level metrics to evaluate rearrangement quality and provide four baselines: a hierarchical task-and-motion planning method, a vision-language-model-based method, and two learning-based approaches (IL and RL). Experimental results demonstrate that current methods struggle to complete the task efficiently, highlighting ESRP as a challenging frontier for embodied agents in scene understanding and long-horizon task planning. This work serves as a stepping stone toward deploying intelligent agents in real-world scenarios.
GRAB-TAMP is introduced, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency.
N. Vijayakumar, Nav Singhal, G. Varma et al.· 0 citations
Planning in complex environments requires task specifications grounded in representations that capture objects, relations, and affordances; scene graphs meet this need, but their size in large environments hinders efficient planning. While task-aware pruning and hierarchical abstractions have been explored, a general,...
Robots are now expected to execute increasingly complex long-horizon tasks in unstructured environments. Despite the strong potential of pretrained Vision-Language Models (VLMs) in task planning, their direct application to robotic manipulation is hindered by logical reasoning deviations and inadequate geometric scene...
Guang-Hui Ma, Jia-Hui Guo, Xin-Hua Tang et al.· Italian National Conference...· 0 citations
Human assistance in robotics spans around several tasks such as navigation, object manipulation, and placement, where a key challenge is selecting target destinations that align with human intentions or preferences. We focus on this challenge in the context of Virtual Placement (VP), the task of identifying all plausib...
Amir Belder, Goncalo Dias Pais, R. Vivanti et al.· 1 citation
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static...
Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang et al.· 0 citations
RoboPhys-3D, a 3D-grounded EWM benchmark built on RoboTwin 2.0, is introduced, and Average Full Score, a hierarchical score averaging all 50 metrics for comprehensive evaluation, and RoboPhyscore, a compact task-aligned score averaging the metrics most strongly correlated with task success are introduced.
Tian-Yi Wang, Jia-Zhou Chen, Yiming Xu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.