Skip to content
Preprint

Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion

Aug 2026 · 0 citations · 56 references
Computer Science

TL;DR

This work proposes Inverted Asymmetric Fusion (IAF), which avoids forcing mutual attention across modalities, and shows that IAF preserves the dominant modality's internal accuracy at its unimodal ceiling across all tested configurations.

Abstract

Fusing multiple modalities is expected to improve model performance. However, on the MultiHuSE dataset, early, late, and symmetric attention fusion often fail to outperform the best unimodal baseline (text). Pathway isolation of a symmetric attention fusion model reveals that the text-pathway accuracy drops from 74.9% to 56.4% after fusion in one such setting, indicating that the dominant modality can be degraded during integration. We term this strong-modality collapse and argue that it helps explain why some multimodal models fail to surpass unimodal baselines. We propose Inverted Asymmetric Fusion (IAF), which avoids forcing mutual attention across modalities. The dominant modality is preserved by passing through fusion unchanged, while weaker modalities attend to it as a contextual anchor. Before fusion, weaker modalities are strengthened using Modality-Aware Knowledge Distillation. We evaluate IAF on three benchmarks with different modality hierarchies: text-dominant datasets (MultiHuSE, UR-FUNNY) and an audio-visual-dominant dataset (MUStARD). Pathway isolation shows that IAF preserves the dominant modality's internal accuracy at its unimodal ceiling across all tested configurations, whereas symmetric fusion degrades it by up to 18.5% on MultiHuSE. IAF improves over the strongest unimodal baseline by up to 8.25%.

View source

Similar papers

Conference Aug 2026

Adaptive Cross-Modal Fusion With Instance-Level Gating for Vision-Language Understanding

Multimodal deep learning integrates heterogeneous data sources such as images and text to enable machines to understand complex real-world contexts. Although recent vision-language models have achieved significant progress, most existing approaches rely on rigid fusion strategies that combine modalities either at early...

Unnati A. Patel, Sanskruti Patel, J. Nanavati et al. · 0 citations
Sep 2026

Understanding Multimodal Learning From Modality Fusion and Alignment Perspectives.

This paper develops a dynamic strategy that jointly optimizes modality fusion and alignment, and develops a learning-based strategy using a bi-level optimization framework and theoretically proves the convergence of the learning algorithm to ensure its reliability.

Yang Yang, Feng-Qiang Wan, Qing-Jun Jiang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

More Perspectives, Stronger Signals: Multi-Perspective Enhancement and Progressive Fusion for Multimodal Entity Representation Learning

PrismF is a unified framework that synergizes multi-perspective enhancement with progressive fusion to extract stronger signals from diverse inputs and improves cross-modal integration through a progressive fusion strategy that dynamically calibrates inter-modal interactions.

Chen-Yi Xiong, Yan Zhang, Jing Hu et al. · 0 citations
#machine learning Preprint Sep 2026

Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification

This paper proposes multimodal Max Confidence Regularization (MaxCR) to dynamically intervene in modality semantic confidence, and reveals that this flaw stems from unimodal characteristics rather than multimodal learning, and this confidence discrepancy can be corrected by positive cross-modal intervention.

Long-Fei Huang, Xiang-Yu Wu, Yang Yang · 0 citations
Open access Sep 2026

Cross-Modal Representation Learning for Integrating Heterogeneous Data in AI Systems

A cross-modal representation learning framework that aligns heterogeneous modalities within a shared latent representation space and exhibits strong robustness under missing modality conditions, with significantly lower performance degradation compared to baseline approaches is proposed.

I. Ibrahim, M. Alshar'e, I. Sanjaya et al. · 0 citations
#natural language process... Preprint Sep 2026

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

Modality robustness under knowledge conflict is studied across 13 MLLMs and two datasets, and it is found that instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks.

Jungyeong Lee, Yejin Yoon, Taeuk Kim · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.