Skip to content
Preprint

ProClosure: Hierarchical Room-Object Assignment using Progressive Boundary Closure from Monocular Video

Sep 2026 · 0 citations · 29 references
Computer Science

TL;DR

Progressive Boundary Closure, which recovers room layer from a monocular RGB video, is introduced, which requires the same treatment: a room should not extend across either, so both are closed and need not be distinguished.

Abstract

A 3D scene graph groups objects into rooms. When a robot is asked to fetch an object from the kitchen, that grouping is what tells it where to look. An object recorded in the wrong room is not retrievable by a query naming the correct room. We introduce Progressive Boundary Closure, which recovers room layer from a monocular RGB video. A SLAM front end and an open-vocabulary segmenter supply a structural point cloud, camera trajectory and object tracks. The cloud is rasterised into a top-down map, rooms are recovered from it, and each object takes the room holding most of its extent. The difficulty lies in the map itself. Walls are recorded only where the camera looked, so a gap in the boundary may be a doorway or a stretch of wall that was never observed; nothing distinguishes the two. Prior methods treat both as passages, merging rooms that should remain separate. We observe that both require the same treatment: a room should not extend across either, so both are closed and need not be distinguished. Such an opening closes under a small amount of boundary growth, and few sightlines cross it, so points in different rooms rarely see one another. We use the first to recover rooms and the second to assign objects to them. Rooms are obtained by Progressively thickening the boundary inward and freezing each free-space region once it becomes enclosed, so every opening seals at its own scale rather than at a radius fixed in advance. Camera poses are used as seeds, which removes the sampling heuristic and makes the segmentation deterministic. Over 10 floors of 6 HM3D-Semantics scenes, scored against HOV-SG on identical top-down maps, we recover 74 rooms for 72 annotated regions (HOV-SG: 44), raising room F_1 from 0.741 to 0.890 at IoU 0.25 at some cost in precision, and object-to-room ARI from 0.488 to 0.696 (p=0.002, ahead on every floor).

View source

Similar papers

Preprint Sep 2026

CoaG: Cylinders on a Grid for Coarse 3D Layout Control in Video Generation

We ask how little geometry a person has to draw to control both where people stand and where the camera moves in a generated video. Our answer is a ground plane and one cylinder per person. A user draws a grid on the ground, places one cylinder where each person should stand, moves the cylinders and the camera over 81...

Zhangsihao Yang, Mengyi Shan · 0 citations
Preprint Sep 2026

Hierarchical Aggregation of Semantic Uncertainty in 3D Scene Graphs

Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a robot querying the map cannot tell which of its entries are unreliable. Estimators of semantic uncertainty could supply that distinction, but they require repeated samplin...

Carlos Roberto Cueto Zumaya, Iacopo Catalano, W. Bessa et al. · 0 citations
#machine learning Preprint Sep 2026

CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image

Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle in such scenes: a missed object is never reconstructed, a merged detection can fuse two objects, and separately reconstru...

Dongwon Son, Junhyek Han, Yoon-Je Cho et al. · 0 citations
Preprint Sep 2026

HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis

Compositional 3D scene reconstruction has recently been explored from two directions: agentic reasoning that provides semantic understanding of spatial relationships but lacks precise alignment with input images; and visual geometry foundation models that predict dense point maps from input images but the reconstructio...

Shuoyao Sun, Chen Wang, En-Xin Song et al. · 1 citation
Preprint Sep 2026

VideoReloc: Long-Term Indoor Video Relocalization against a Kilobyte-Scale Semantic Scene Graph

This work introduces VideoReloc, whose adaptive clips use odometry to gather spatial evidence until object and motion criteria are met, adapting query length to the observed scene, and reframes sparse-map relocalization as verification of spatially extended video queries.

Qian-Ru Li, Xu-Yang Chen, Xu-Qin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.