A multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion is adopted, which considers that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario the authors refer to as missing arbitrary modalities.
Abstract
Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle missing modalities, prior studies have mainly covered bimodal data setups and focused on designing robust fusion processes. Instead, we adopt a multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion. Specifically, we consider that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario we refer to as missing arbitrary modalities. To address this challenge, we introduce two alternative approaches that leverage information at both feature- and decision-level. Experiments on two multi-modal classification benchmarks demonstrate significant robustness gains in various missing modality conditions. The first method shows more robust behavior under minimal missing conditions, where a single modality is absent, whereas the second performs better under extreme missing conditions, where all-but-one modalities are missing. Our code is available at https://github.com/fmenat/Co4Miss.
A cross-modal representation learning framework that aligns heterogeneous modalities within a shared latent representation space and exhibits strong robustness under missing modality conditions, with significantly lower performance degradation compared to baseline approaches is proposed.
I. Ibrahim, M. Alshar'e, I. Sanjaya et al.· Journal of Data Science· 0 citations
Modality-Conditioned Conformal Fusion is, to the authors' knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty...
The T2ID is proposed, a multimodal diagnostic framework, which adaptively completes missing modalities and explicitly models cross-modal interactions while maintaining trustworthiness and robustness against missing modality scenarios.
Jing Li, Qin-Kai Yu, Fei-Xiang Zhou et al.· Medical Image Analysis· 0 citations
Flux is proposed, a multimodal federated learning framework built around two complementary components, modality-aware confidence tempering and gradient-decoupled private adaptation, that enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared represe...
Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin et al.· 0 citations
Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them fo...
Zirui Cheng, Xun Xu, Tiankai Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.