Skip to content
Open access

Object-Centric 2D-to-3D Pipeline for Interior-Design Visualization: Reference-Free Asset Evaluation and a Structured3D Scene-Level Benchmark

Jul 2026 · Electronics · 0 citations · 13 references

Abstract

This study presents a modular AI-assisted workflow for converting single 2D interior images into textured 3D assets and for evaluating those assets when ground-truth 3D meshes are unavailable. The proposed pipeline combines object detection, instance isolation, monocular-depth estimation, image-to-3D generation, texture synthesis, mesh export, and cloud-based execution to support early-stage interior-design and real-estate visualization tasks. A reference-free validation protocol is introduced, based on rendered multi-view comparisons, silhouette Intersection-over-Union, automated captioning, and multimodal embedding similarity, and is complemented by a composite validation framework that benchmarks reconstructed scenes against 200 panoramic indoor scenes from the Structured3D dataset using Hungarian-matched placement, size, recall, and relative-distance metrics. The workflow was implemented and tested using contemporary computer-vision and generative 3D components, with Hunyuan3D 2.0 used as the main reconstruction model. Proof-of-concept experiments on a representative corpus of 178 synthetically generated single-object images spanning a range of interior furniture categories show comparable silhouette IoU for textured and non-textured outputs and indicate that texture-preserving renderings improve visual and semantic similarity scores across CLIP-based evaluations. The 200-scene dataset evaluation reveals stable spatial localization (placement error ≈ 1.18 m, relative-distance error ≈ 0.54 m) alongside systematic over-prediction and size-calibration errors. Beyond the applied pipeline, the study contributes a reference-free, ground-truth-free protocol for 3D-asset evaluation and a first quantified account of where object-centric single-image reconstruction is reliable—spatial placement—and where it is not—object scale and spurious detection—at interior-scene scale. The results demonstrate the feasibility of integrating perception, 3D reconstruction, semantic assessment, and scalable deployment into a single applied pipeline, while remaining proof-of-concept and requiring extension to larger object and scene corpora, baselines, real-photograph evaluation, and human-centered assessment before broad claims about general interior-scene reconstruction can be made.

Read PDF