Skip to content

Apples on the Table? Evaluating Text-Guided 3D Scene Synthesis via Fine-Grained Constraint Verification

Nov 2025 · 0 citations · 44 references
Computer Science

TL;DR

LEGO, a benchmark dataset that pairs each user description with human-annotated constraints and a reference scene, and LEGO-Eval, an evaluation framework that decomposes a description into atomic constraints and verifies each one using tools that ground textual references to 3D objects and reason about their spatial relationships are introduced.

Abstract

Accurately synthesizing 3D scenes from user-provided text descriptions is crucial for developing embodied agents. Despite the importance of scene-description alignment, existing evaluation methods for such text-guided 3D scene synthesis either capture only coarse similarity between the synthesized scene and the user description, or ignore the spatial reasoning for verifying object placement. None of them addressed the fine-grained constraints (e.g., X needs to be in the scene in a Y manner) implied by the description from users. To address this, we introduce LEGO, a benchmark dataset that pairs each user description with human-annotated constraints and a reference scene, and LEGO-Eval, an evaluation framework that decomposes a description into atomic constraints and verifies each one using tools that ground textual references to 3D objects and reason about their spatial relationships. We show that (i) LEGO-Eval evaluates misalignment far more accurately than existing methods and (ii) current scene synthesis approaches achieve at most 10% success rate in LEGO-Eval.

View source

Similar papers

Preprint Aug 2026

ReRef-3D: A Benchmark for Spatial Referring Expression-Guided 3D Scene Rearrangement

We introduce ReRef-3D, a benchmark for language-guided placement in 3D scenes. It contains 33,826 instructions across 998 CLEVR-derived scenes, spanning 16 placement families and direct, one-hop, and two-hop references. Each instruction must be resolved into a valid new placement position. Given that an instruction def...

Mary Lynn Martin, Yifei Zhang, Martha Palmer et al. · 0 citations
Preprint Aug 2026

DesignAgent3D: Interactive 3D Scene Editing via Designer-like Multimodal Reasoning

DesignAgent3D is presented, an interactive multimodal agentic framework that reformulates 3D scene editing as a designer-like Plan-Perceive-Act paradigm, delivering superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency.

Xiu-Jin Liu, Tian-Yu Yang, Yilun Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
Preprint Aug 2026

CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels

3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual. Existing language-guided approaches in 3D scenes mainly focus on object-centric grounding, determining what to observe but rarely controlling how it should...

Ji-Rong Li, Satoshi Ikehata, Shuhei Kurita et al. · 0 citations
Preprint Sep 2026

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstructio...

Shuoyao Sun, Chen Wang, En-Xin Song et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.