RoomWright is presented, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction, providing interactive environments for embodied AI and policy learning.
Abstract
Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments, yet they remain focused on visual construction and object-level articulation, leaving the functional usage of scenes largely unmodeled. To address this problem, we present RoomWright, an agentic usage-driven framework for generating 3D scenes represented entirely as code for embodied interaction. RoomWright performs usage-driven object reasoning, which treats each anchor as a task centre and admits task-required objects and their affordances. A code agent further enables multi-part interaction by compiling each interaction into a trigger, condition, effect rule that updates structured object states, capturing causal dependencies across objects. Moreover, since manipuland orientation is ambiguous and hard to recover from pixels, RoomWright alleviates this via annotation-informed usage-guided orientation. Extensive experiments demonstrate the effectiveness of our method. The resulting scenes are executable, editable, and simulation-ready, providing interactive environments for embodied AI and policy learning.
Embodied agents need environments that are visually diverse, physically interactive, and changing over time. Procedural simulators can generate large interactive scene collections, and recent 4D generators produce compelling visual dynamics. Combining these properties in one environment, however, still demands extensiv...
Ze-Hao Qi, Hao-Chen Luo, Jia-Wang Bian et al.· 0 citations
This work presents RoboInter1.5, an extended and holistic suite of intermediate representations for both robotic manipulation and embodied world modeling, and introduces RoboInter-World, which leverages intermediate representations as structured conditioning signals for controllable prediction of future world states.
Ziqin Wang, Hao Li, Weijun Wang et al.· arXiv.org· 0 citations
Learning-based manipulation requires supervision that is both semantically meaningful and physically executable, but current data pipelines often provide only one of these properties. Human demonstrations capture intent but are costly to collect and constrained by the human-robot embodiment gap, while simulation can sc...
Yian Wang, Jun-Yi Cao, Xiao-Wen Qiu et al.· 0 citations
DesignAgent3D is presented, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.
Xiu-Jin Liu, Tian-Yu Yang, Yilun Zhao et al.· 0 citations
Function-Room Generation is introduced, a new indoor 3D scene generation setting that creates rooms supporting explicit functional goals rather than merely visually plausible layouts, and ScenePRM, an execution-grounded process reward framework that improves the expert through reinforcement learning with functional, ge...
Hao Feng, Zhi Zuo, Ming-Jian Liang et al.· 1 citation
Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.
Jia Luo· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.