Skip to content
Preprint

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

Aug 2026 · 1 citation · 71 references
Computer Science

TL;DR

iARCS is presented, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements and shows that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.

Abstract

Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements. iARCS uses a two-phase strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific finetuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.

View source

Similar papers

Preprint Sep 2026

GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning

GIF, an agentic Generation framework for Interactive and Functional object compositions, recast this problem as disentangled reconstruction followed by relative pose recovery, revealing diversity scaling in both simulation and real-world deployment.

Long-Ji Xu, Zhi-Qi Zhang, Mi Yan et al. · 2 citations · ⚡2
Preprint Sep 2026

ARSTAG: An Agentic Real2Sim2Real System for Task-Specific Robot Data Generation

ARSTAG is presented, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data and shows that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size.

Bo-Wei Li, Yun-Er Zhang, Chang-Liu Liu · 0 citations
Preprint Sep 2026

SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placeme...

Xingjian Ran · 0 citations
Preprint Aug 2026

4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied Simulation

4DSynth is presented, a controllable procedural system that turns a natural-language description, a blueprint mask, or a single photograph into an editable 4D environment with explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation state.

Ze-Hao Qi, Hao-Chen Luo, Jia-Wang Bian et al. · 0 citations
Preprint Sep 2026

MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence

MachEmbodied-U0, a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture, is presented, a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture that combines competitive...

Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.