Modality-Conditioned Conformal Fusion is, to the authors' knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty to the absent modality responsible.
Abstract
Multimodal fusion architectures typically assume all modalities are available at inference, yet sensor failures, acquisition variability, and cost constraints routinely produce incomplete observations. Existing work treats modality absence as a prediction-accuracy problem, leaving a more basic question unanswered: whether a model's confidence estimates remain calibrated when an entire input stream is removed. We argue that missing-modality robustness and calibrated uncertainty are a single coupled property, and introduce Modality-Conditioned Conformal Fusion (MCCF), an architecture that addresses both at once. MCCF combines a multimodal bottleneck fusion backbone trained with modality dropout, per-modality evidential heads producing modality-decomposed Dirichlet distributions, and a Dempster-Shafer combination rule that fuses the per-modality evidence into a joint predictive distribution; an absent modality contributes vacuous evidence that is structurally ignored, so the fused uncertainty automatically reflects the reduced information without test-time imputation. A Mondrian conformal calibration module keyed on the modality-presence mask then provides finite-sample group-conditional coverage for every non-empty modality subset. MCCF is, to our knowledge, the first method with formal coverage guarantees under arbitrary modality availability through architectural integration rather than post-hoc recalibration, and the evidential decomposition yields per-modality vacuity scores that localise uncertainty to the absent modality responsible. Across a synthetic problem and three real multimodal benchmarks, MCCF holds its target coverage on every modality-presence subset, substantially narrows the coverage gap between full and partial modalities relative to a marginal split-conformal baseline, and imposes no measurable accuracy cost relative to temperature-scaled and evidential baselines.
RiVaT-Fuse is proposed, a reliability-calibrated variational tensor fusion framework that defines fusion as sample-wise latent-state estimation and achieves the strongest overall predictive rank among direct representation-level baselines while improving probability and label stability under perturbation.
Experiments on multimodal intent-recognition benchmarks show that PRIME maintains competitive clean-data performance while improving robustness under missing, noisy, conflicting, and modality-imbalanced conditions.
Suraj Kumar, Mohnish Raj, Soumi Chattopadhayay et al.· 0 citations
A modality-aware conformal calibration layer that trains or reuses one predictor per modality, computes a disagreement score from their predictions, and uses that score in split conformal calibration under a strict split protocol for simple, model-agnostic reliability layer for multi-modal regression systems.
GAUGE is proposed, a lightweight counterfactual gating framework for incomplete multimodal classification that outperforms strong baselines across diverse incomplete-input settings and is established as a principled and scalable framework for fine-grained evidence control under modality incompleteness.
Yun Shi, Enshui Yu, Kai-Rui Guo et al.· 0 citations
A unified framework that formulates missing-modality generation as a linear inverse problem under a joint distribution and solves it via posterior sampling with a flow matching model is proposed, which can reconstruct arbitrary missing modalities at inference time by guiding the sampling trajectory to enforce measureme...
A multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion is adopted, which considers that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario the authors refer to as missing arbitrary modalities.
F. Mena, Dino Ienco, R. Interdonato et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.