Skip to content
Preprint

Multimodal Floorplan Encoding: Learning Dense Modality-Invariant Representations

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

The Multimodal Floorplan Encoder (MMFE) is introduced, which maps diverse 2D indoor representations into a shared dense latent grid and improves cross-modal dense matching, enables robust similarity alignment with RANSAC, and yields strong retrieval when paired with learned aggregation.

Abstract

Floorplans arise in many forms, from vector CAD drawings to raster renderings and sensor-derived density maps. This heterogeneity makes it difficult to build learning systems that transfer across modalities and support geometry-centric tasks such as alignment and retrieval. We introduce the Multimodal Floorplan Encoder (MMFE), which maps diverse 2D indoor representations into a shared dense latent grid. MMFE combines a frozen DINOv3 backbone with a trainable Dense Prediction Transformer (DPT) head, and is trained with a per-cell Information Noise-Contrastive Estimation (InfoNCE) objective that aligns spatially corresponding regions across modalities while using all other cells as negatives. To improve robustness to geometric distortions, we incorporate controlled similarity transformations and enforce geometric consistency through feature-grid warping. On Structured3D, a held-out out-of-domain dataset, MMFE improves cross-modal dense matching, enables robust similarity alignment with RANSAC, and yields strong retrieval when paired with learned aggregation.

View source

Similar papers

Preprint Sep 2026

From Alignment to Fusion in 3D Vision-Language

Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous represe...

Xue-Qi Qiu, Xing-Yu Miao, Jing-Jing Deng et al. · 0 citations
Preprint Aug 2026

PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching

Accurate stereo matching remains challenging in ill-posed regions such as fine structures, reflective, or transparent objects, where appearance cues are often ambiguous or unreliable. To tackle this, we propose PhasorNet, a lightweight yet powerful framework that boosts geometric discrimination via frequency-domain cue...

Md Raqib Khan, S. Vipparthi, Subrahmanyam Murala · 0 citations
Preprint Oct 2026

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partia...

Akira-Miranda Adeyomi Adeniran-Lowe, B. Singh, Lars Arnold Dethlefsen et al. · 0 citations
Preprint Sep 2026

Scene Retargeting: Learning Object Placement with Analogical Transfer

Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We f...

Minkwan Kim, Junho Kim, Seung-Min Lee et al. · 0 citations
Preprint Aug 2026

TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching

Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods, existing approaches primarily optimize matching within a single model and often struggle to maintain reliable correspondences under challenging conditions such as noisy...

Chong Wang, Jun-Jie Gao · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.