CM-MAE is presented, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer that builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives.
Abstract
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.
A physics-guided simulator that synthesizes orientation-diverse wireless training data from single-orientation motion is developed and a dual-attention network that extracts activity-discriminative and orientation-robust representations from dual-link Doppler spectrograms is proposed.
Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment. In this paper, we propose BeyondFusion,...
Minchong Chen, Xiaoyun Yuan, Minyu Cao et al.· arXiv.org· 0 citations
It is shown that removing image reconstruction relocates the central design problem to the lift: how 1D measurements become a 2D task representation, and measurement-to-space adaptivity organizes both the clean-to-noisy operating envelope and the failure a system encounters.
In optical measurement environments for weapons system test and evaluation, imagery is frequently degraded by haze and smoke, impairing downstream analysis. Existing methods rely on physics-based models or single-image deep learning, both struggling under non-uniform haze or recovering occluded structures. This paper p...
Sangin Lee· Journal of the Korea Institu...· 0 citations
Cyclops is proposed, a framework that translates sparse Non-Repetitive Scanning LiDAR intensity into RGB video, enabling camera-free inference for all-day perception tasks and mitigating inter-frame flickering.
Wei Gao, Jian Shu, Ming-Le Zhao et al.· 0 citations
Side-Scan Sonar (SSS) is a primary modality for large-scale underwater mapping, yet automated perception and cross-modal alignment are severely bottlenecked by acoustic complexities such as speckle noise, shadows, and extreme viewpoint dependencies. Traditional handcrafted descriptors and modern deep learning matchers...
Taqi Hamoda, Nuno Gracias· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.