The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation, supported by Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66.
Abstract
Optical imagery provides rich appearance cues, whereas synthetic aperture radar (SAR) offers observations that are less sensitive to illumination and weather, making optical--SAR fusion attractive for remote-sensing object detection. However, the presence of multiple modalities does not guarantee beneficial fusion: imperfect spatial, temporal, and semantic correspondence can make an otherwise intact stream conditionally harmful and induce negative cross-modal transfer. We handle this issue through a model-specific task-utility perspective and learn task-conditioned contribution routing using detection supervision alone. The proposed fusion-boundary-aligned routing regulates each modality's contribution before the first learned cross-modal feature-value mixing operation. For architectures with frequent shallow interaction, a Feature Router performs cross-conditioned, group-addressable modulation near the input; for dual-backbone architectures, a Dual-Statistic Semantic Router predicts stream-level contribution weights from modality-specific average and maximum statistics before late semantic fusion. The routers require no explicit utility supervision, quality labels, reconstruction, or distillation. Experiments on M4-SAR and SpaceNet6-OTD cover nominal full inputs, controlled correspondence shifts, missing modalities, and four nonzero modality-corruption scenarios. Across the reported clean-training controls, routing improves full-input $\text{mAP}_{50}$ by 0.5--5.9 points. Relative to the corresponding modality-dropout baselines, it raises missing-modality $\text{mAP}_{50}$ by 7.6--41.6 points and reduces the negative-transfer rate by up to 12.7 percentage points. Spearman correlations between the learned routing weights and model-specific leave-one-modality-out utility range from 0.45 to 0.66, supporting the task-utility interpretation of the routing coefficients.
Optical imagery provides rich spectral and texture cues but is vulnerable to cloud cover and imaging conditions, whereas synthetic aperture radar (SAR) offers all-weather observation but contains speckle noise and geometry-dependent distortions. Existing optical–SAR segmentation methods often treat local spatial correc...
Hao-Tian Liu· International Conference on...· 0 citations
Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make...
A query-based multimodal fusion framework, termed SRCDet, is proposed for camera-4D radar fusion, which achieves consistent improvements across nearly all metrics and low error rates in clear and adverse weather conditions, highlighting its practical adaptability to automotive-grade systems and effectiveness in safety-...
Wen-Jin Ai, Lianqing Zheng, Long Yang et al.· Measurement science and tech...· 0 citations
M-FSAD-KD is proposed, a full-link distillation framework whose neck-stage Fourier-gated alignment transfers low-frequency structural content while preserving target-edge high-frequency content; a joint spatial–channel attention mask, a shallow backbone adapter, and a response-level knowledge distillation (KD) term com...
Yu-Ming Tong, Kai-Na Xiong, Jun Liu et al.· Remote Sensing· 0 citations