iARCS is presented, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements and shows that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
Abstract
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements. iARCS uses a two-phase strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific finetuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
GIF, an agentic Generation framework for Interactive and Functional object compositions, recast this problem as disentangled reconstruction followed by relative pose recovery, revealing diversity scaling in both simulation and real-world deployment.
Long-Ji Xu, Zhi-Qi Zhang, Mi Yan et al.· 2 citations· ⚡2
ARSTAG is presented, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data and shows that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size.
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet their scalable generation remains challenging. Recent agentic text-to-3D scene pipelines that rely on vision-language models (VLMs) can generate scenes of high fidelity but require costly iterative object placeme...
4DSynth is presented, a controllable procedural system that turns a natural-language description, a blueprint mask, or a single photograph into an editable 4D environment with explicit geometry, animated actors, collision-free trajectories, and physics-ready simulation state.
Ze-Hao Qi, Hao-Chen Luo, Jia-Wang Bian et al.· 0 citations
MachEmbodied-U0, a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture, is presented, a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture that combines competitive...
Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.