Skip to content
Preprint

CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications

Aug 2026 · 0 citations · 19 references
Computer Science

TL;DR

CM-MAE is presented, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer that builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives.

Abstract

Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.

View source

Similar papers

Preprint Aug 2026

Dual-Attention and Adversarial Transfer Networks for Sim-to-Real Cross-Orientation Wireless Sensing

A physics-guided simulator that synthesizes orientation-diverse wireless training data from single-orientation motion is developed and a dual-attention network that extracts activity-discriminative and orientation-robust representations from dual-link Doppler spectrograms is proposed.

L. Du, Kehan Wu, Tong Zhang et al. · 0 citations
Jul 2026

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion,...

Minchong Chen, Xiaoyun Yuan, Minyu Cao et al. · 0 citations
Jul 2026

The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing

It is shown that removing image reconstruction relocates the central design problem to the lift: how 1D measurements become a 2D task representation, and measurement-to-space adaptivity organizes both the clean-to-noisy operating envelope and the failure a system encounters.

Yu-Yuan Han, Jing-Wei Li, Xiaoxia Zhang et al. · 0 citations
Open access Aug 2026

Reference-Guided Global-Local Context Fusion Network for Residual Image Restoration

In optical measurement environments for weapons system test and evaluation, imagery is frequently degraded by haze and smoke, impairing downstream analysis. Existing methods rely on physics-based models or single-image deep learning, both struggling under non-uniform haze or recovering occluded structures. This paper p...

Sangin Lee · 0 citations
Preprint Aug 2026

Cyclops: LiDAR as a Camera That Dreams in Color

Cyclops is proposed, a framework that translates sparse Non-Repetitive Scanning LiDAR intensity into RGB video, enabling camera-free inference for all-day perception tasks and mitigating inter-frame flickering.

Wei Gao, Jian Shu, Ming-Le Zhao et al. · 0 citations
Preprint Aug 2026

Geometry-Driven Opti-Acoustic Co-Registration and View-Invariant Reflectivity Mapping for Side-Scan Sonar

Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlenecked by acoustic complexities such as speckle noise, shadows, and extreme viewpoint dependencies. Traditional handcrafted descriptors and modern deep learning matchers...

Taqi Hamoda, Nuno Gracias · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.