Skip to content
Preprint Open access

PolyLayout: Multi-room Manhattan Layout Estimation

Aug 2026 · 0 citations · 52 references
Computer Science

TL;DR

This work proposes PolyLayout, a multi-room layout estimation method that parameterizes room layouts as Manhattan 3D polygons and optimizes them jointly across multiple rooms, and introduces two new multi-view multi-room layout benchmarks by providing layout annotations to existing datasets.

Abstract

Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poor generalization to new datasets or restrictive geometric assumptions of the room shape or camera configuration. Most also estimate rooms independently, failing to exploit shared building structure such as dominant directions, ground plane or ceiling height. We propose PolyLayout, a multi-room layout estimation method that parameterizes room layouts as Manhattan 3D polygons and optimizes them jointly across multiple rooms. The optimization objective is predicted by a neural network on top of robust pre-trained visual features and trained end-to-end with supervision only on output room layouts. At the same time, camera projection and polygon updates remain explicit and model-based. This separation between learned scoring and geometry improves generalization to new datasets and camera parameters. During optimization, PolyLayout adaptively refines the polygon topology through iterative wall split and merge operations while jointly utilizing structural cues across rooms. We introduce two new multi-view multi-room layout benchmarks by providing layout annotations to existing datasets, and experiments show that PolyLayout outperforms prior approaches, both in terms of accuracy and robustness. Project page: https://ghanning.github.io/PolyLayout

Read PDF

Similar papers

Preprint Aug 2026

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Prediction

OccAnyScene is proposed, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model which employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple...

Junjie Liu, Wan-Shui Gan, Zi-Tong Dai et al. · 0 citations
Preprint Sep 2026

GALoc: Gravity Aligned Wireframes for Depth-Free Monocular Floorplan Localization

GALoc, a geometry-first framework that replaces depth prediction with gravity-aligned wireframes that satisfy verticality and coplanarity by construction, is proposed and evaluated end-to-end on Structured3D, with calibrated noise on Gibson, and on real-world author-collected sequences.

Jeahn Han, Minji Kim, Jeongbin Sohn et al. · 0 citations
Preprint Sep 2026

Multimodal Floorplan Encoding: Learning Dense Modality-Invariant Representations

The Multimodal Floorplan Encoder (MMFE) is introduced, which maps diverse 2D indoor representations into a shared dense latent grid and improves cross-modal dense matching, enables robust similarity alignment with RANSAC, and yields strong retrieval when paired with learned aggregation.

X. Anadón, Rémi Pautrat, Rui Wang · 0 citations
Preprint Aug 2026

UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

UniQuery4R is presented, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention, and introduces a direction-magnitude parameterization of scene flow with separate super...

Tiancheng Chen, Sheng Tang, Wenhua Jin et al. · 1 citation
Sep 2026

Unified feature Gaussian representation for robust camera relocalization

LoFG is presented, a localization-oriented Feature Gaussian representation that unifies both stages within a single Gaussian scene and improves the robustness of sparse initialization and the accuracy of dense refinement, demonstrating potential for localization applications in AR, robotics, and visual navigation syste...

Zhen-Dong Xiao, Zi-Ling Wen, Jun Yin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.