Counterfactual Modality Attribution (CMA) is proposed, the first framework for quantifying modality-level contributions in MLLMs, providing a principled framework for auditing multimodal foundation models in safety-critical applications.
Abstract
Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. While existing explainability methods identify influential image regions or text tokens, they cannot answer a fundamental question: which modality drives a prediction? Consequently, a model may produce the correct output while relying on the wrong source of evidence, masking shortcut learning and unsafe reasoning. We formulate modality attribution as a complementary explainability objective for multimodal foundation models and propose Counterfactual Modality Attribution (CMA), the first framework for quantifying modality-level contributions in MLLMs. CMA generates image-only, text-only, and joint multimodal counterfactuals using coupled diffusion priors and converts them into principled modality attribution scores through a cooperative game-theoretic formulation based on Shapley values. We evaluate CMA on controlled synthetic benchmarks with known ground-truth modality reliance and on a real-world multimodal clinical dataset. CMA correctly identifies the decision-driving modality in 98% of controlled cases and consistently outperforms baselines, revealing failures of cross-modal reasoning that remain invisible to predictive accuracy alone. Our results establish modality attribution as a complementary dimension of explainability beyond feature attribution, providing a principled framework for auditing multimodal foundation models in safety-critical applications.
GAUGE is proposed, a lightweight counterfactual gating framework for incomplete multimodal classification that outperforms strong baselines across diverse incomplete-input settings and is established as a principled and scalable framework for fine-grained evidence control under modality incompleteness.
Yun Shi, Enshui Yu, Kai-Rui Guo et al.· 0 citations
Modality robustness under knowledge conflict is studied across 13 MLLMs and two datasets, and it is found that instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks.
This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action.
Modality Subspace Activation (MSA) is proposed, a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths and dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.
Hongbo Jiang, Jie Li, Yunhang Shen et al.· 0 citations
HalluPrism, a behavioral diagnostic that re-runs an answer after visual degradation, blank-image replacement, and grounding or relation checks is proposed, a behavioral diagnostic that separates failure diagnosis from abstention scoring.
Aman Prakash, Sourish Dasgupta, Tanmoy Chakraborty· 0 citations
OPD-V is introduced, a visual OPSD paradigm that instantiates privileged information through the Positive Teacher and Negative Teacher that consistently improves reasoning performance while reducing training cost.
Aniri, Jinhe Bi, Peng Liao et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.