This work proposes Inverted Asymmetric Fusion (IAF), which avoids forcing mutual attention across modalities, and shows that IAF preserves the dominant modality's internal accuracy at its unimodal ceiling across all tested configurations.
Abstract
Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9% to 56.4% after fusion in one such setting, indicating that the dominant modality can be degraded during integration. We term this strong-modality collapse and argue that it helps explain why some multimodal models fail to surpass unimodal baselines. We propose Inverted Asymmetric Fusion (IAF), which avoids forcing mutual attention across modalities. The dominant modality is preserved by passing through fusion unchanged, while weaker modalities attend to it as a contextual anchor. Before fusion, weaker modalities are strengthened using Modality-Aware Knowledge Distillation. We evaluate IAF on three benchmarks with different modality hierarchies: text-dominant datasets (MultiHuSE, UR-FUNNY) and an audio-visual-dominant dataset (MUStARD). Pathway isolation shows that IAF preserves the dominant modality's internal accuracy at its unimodal ceiling across all tested configurations, whereas symmetric fusion degrades it by up to 18.5% on MultiHuSE. IAF improves over the strongest unimodal baseline by up to 8.25%.
Multimodal deep learning integrates heterogeneous data sources such as images and text to enable machines to understand complex real-world contexts. Although recent vision-language models have achieved significant progress, most existing approaches rely on rigid fusion strategies that combine modalities either at early...
Unnati A. Patel, Sanskruti Patel, J. Nanavati et al.· International Conference on...· 0 citations
This paper develops a dynamic strategy that jointly optimizes modality fusion and alignment, and develops a learning-based strategy using a bi-level optimization framework and theoretically proves the convergence of the learning algorithm to ensure its reliability.
Yang Yang, Feng-Qiang Wan, Qing-Jun Jiang et al.· IEEE Transactions on Pattern...· 0 citations
PrismF is a unified framework that synergizes multi-perspective enhancement with progressive fusion to extract stronger signals from diverse inputs and improves cross-modal integration through a progressive fusion strategy that dynamically calibrates inter-modal interactions.
Chen-Yi Xiong, Yan Zhang, Jing Hu et al.· 0 citations
This paper proposes multimodal Max Confidence Regularization (MaxCR) to dynamically intervene in modality semantic confidence, and reveals that this flaw stems from unimodal characteristics rather than multimodal learning, and this confidence discrepancy can be corrected by positive cross-modal intervention.
Long-Fei Huang, Xiang-Yu Wu, Yang Yang· 0 citations
A cross-modal representation learning framework that aligns heterogeneous modalities within a shared latent representation space and exhibits strong robustness under missing modality conditions, with significantly lower performance degradation compared to baseline approaches is proposed.
I. Ibrahim, M. Alshar'e, I. Sanjaya et al.· Journal of Data Science· 0 citations
Modality robustness under knowledge conflict is studied across 13 MLLMs and two datasets, and it is found that instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks.
Jungyeong Lee, Yejin Yoon, Taeuk Kim· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.