Skip to content

OpenAsset: A Pipeline for Open-World Asset Integration into Indoor Scene Synthesis

Sep 2026 · Electronics · 0 citations · 15 references
3D Shape Modeling and Analysis

TL;DR

This paper proposes OpenAsset, a novel framework that converts single images of real-world objects into reusable, canonicalized 3D assets that seamlessly integrate into the scene synthesis workflow and introduces an automated canonicalization module that normalizes the scale, orientation, and coordinate systems of the reconstructed meshes.

Abstract

Text-driven 3D indoor scene synthesis has witnessed significant progress through Large Language Model (LLM)-based frameworks like ReSpace. However, current paradigms heavily rely on retrieving objects from pre-defined, static 3D asset libraries, which fundamentally constrains the diversity and personalization of generated scenes due to the closed-set nature of existing databases. Conversely, recent breakthroughs in promptable segmentation (SAM 3) and single-image 3D reconstruction (SAM 3D) have empowered the extraction of high-fidelity 3D geometry and texture from in-the-wild images. In this paper, we bridge the gap between text-driven scene layout generation and single-view object reconstruction by proposing OpenAsset. This novel framework converts single images of real-world objects into reusable, canonicalized 3D assets that seamlessly integrate into the scene synthesis workflow. Specifically, given a user concept prompt or target region, OpenAsset leverages SAM3 for precise instance isolation and SAM3D for geometry and texture recovery. To ensure compatibility with structured scene representations (SSR), we introduce an automated canonicalization module that normalizes the scale, orientation, and coordinate systems of the reconstructed meshes. By transforming “wild” visual percepts into standardized assets, OpenAsset effectively expands the controllable vocabulary of indoor scene synthesis beyond curated datasets, offering a practical pathway toward user-sourced open 3D scene generation that supports custom objects outside fixed predefined asset libraries. It is worth noting that the satisfactory performance of OpenAsset critically depends on effective segmentation and reconstruction results, which serve as essential prerequisites for our method.

Read PDF

Similar papers

Preprint Aug 2026

Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding

A comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems, and an innovative manifold-distance-based point cloud refinement strategy.

Jie Xu, Na Zhao · 0 citations
Preprint Aug 2026

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

SPAR is proposed, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation that structurally couples motion estimation with multi-view visual and semantic learning and reveals a strong inter-task synergy between photometric scene reconstru...

Boyu Cai, Li Yang, Yan Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
Conference Open access Sep 2026

Interactive Open-Set Semantic Mapping with a 3D Scene Graph Backend

A modular mapping architecture is demonstrated that establishes 3D Semantic Scene Graphs (3DSSGs) as its foundational back-end, enabling the dense representation of extensive environments containing thousands of unique object instances and supporting open-vocabulary queries via CLIP features without requiring any addit...

Felix Igelbrink, Lennart Niecksch, Martin Günther et al. · 0 citations
#artificial intelligence Preprint Oct 2026

RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation

Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation...

Minsu Kim, Jaesung Choe, J. Lee et al. · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.