Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 18 references
Abstract
In radiology, multimodal vision-language models (VLMs) are increasingly used for clinical tasks such as report generation and visual question answering. Their adoption has raised important concerns regarding the transparency and trust-worthiness of the generated text reports, as the reasoning process leading to a given report remains opaque for clinicians. In this work, we investigate how post-hoc attribution methods behave in large generative volumetric VLMs that jointly process 3D scans and clinical text, a setting that remains largely unexplored. We propose to use Post-hoc gradient-based attribution that directly links small changes in the input volume to changes in the model's output probability. The faithfulness and spatial specificity of four attribution methods are subsequently characterized and compared when applied to Med3DVLM, a recent medical VLM for image-text understanding. To enable differentiable attribution at inference time, we apply teacher forcing exclusively during the attribution forward pass so that the stochastic generation path is converted into a differentiable computation graph. The cumulative log-likelihood of the generated response is used as the scalar attribution target. Input-level methods, including Saliency and Integrated Gradients, alongside feature-level techniques such as Grad-CAM and Guided Grad-CAM, are computed on volumetric inputs across diverse clinical question types. Qualitative assessment and a quantitative voxel deletion protocol indicate that input-space gradient methods, especially Integrated Gradients, produce spatially selective and causally faithful relevance maps, whereas feature-level methods like Layer Grad-CAM exhibit significant spatial diffusion and fail to provide discriminative localization in multimodal architectures. These results demonstrate that input-level attributions exhibit a significantly steeper drop in model confidence compared to feature-level methods, confirming they are more faithful and spatially precise, while intermediate-layer maps remain diffuse and non-discriminative.
KANEx is introduced, the first ever framework that leverages the symbolic transparency of KANs to ground VLM reasoning, and suggests that grounding linguistic explanations and visual attributions in mathematically interpretable units is a necessary step toward trustworthy medical AI.
Krithi Shailya, Ananya Lakshmi Ravi, V. VenkatanathanK. et al.· arXiv.org· 0 citations
A large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm that unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training is introduced.
The results indicate that confidence-guided masked diffusion with reliability-aware decoding is a useful direction for controllable and reliability-aware clinical assistants.
Saqib Qamar, Goram Mufarah M. Alshmrani· Technologies· 0 citations
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-...
Bo Xu, Quan-Hao Zhu, Bo-Lin Zhu et al.· Proceedings of the Thirty-Fi...· 1 citation
A clinically curated Pan-Asia WSI--report dataset is introduced and the REG 2025 benchmark is established as a benchmark for evaluating WSI-based structured report generation and vision-language understanding in computational pathology, providing insights for the design of clinically grounded multimodal pathology model...
Yu-Mi Lee, Harim Oh, Hyo-yun Kim et al.· 0 citations
Medical report generation aims to reduce the reporting burden of radiologists by translating medical images into clinically meaningful descriptions. Although generative adversarial networks (GANs) have shown potential for image captioning and report generation, their application to radiology reports remains challenging...
Yuan Wang, Shijie Xu, Kun Zhou et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.