Skip to content
Conference

MedEvidence-RAG for Joint Evidence Retrieval and Generator Alignment in Medical Visual Question Answering and Report Generation

Aug 2026 · 2026 3rd International Conference on Image Processing, Multimedia Technology and Machine Learning (IPMML) · pp. 71-75 · 0 citations · 13 references

Abstract

Medical vision-language models remain vulnerable to unsupported findings because their outputs are weakly grounded in images and external clinical evidence. We propose MedEvidence-RAG, which couples a mixture-of-experts (MoE) multimodal retriever with retrieval-conditioned generator alignment. The retriever uses a SigLIP-style objective to match multimodal image-question queries with reports, while supervised fine-tuning (SFT) and direct preference optimization (DPO) teach the generator to use the same retrieved context available at inference. Across IU-Xray, MIMIC-CXR, and Quilt-1M, MedEvidence-RAG improves medical visual question answering (VQA), reaching 89.98 F1 and 91.29 area under the receiver operating characteristic curve (AUC) on MIMIC-CXR. It also achieves the best ROUGE-L on both report-generation benchmarks. Ablations show that retrieval and alignment provide complementary gains, and retrieval analysis confirms more relevant evidence than a Contrastive Language-Image Pre-training (CLIP)-based retriever. These results support joint optimization of evidence acquisition and evidence use for reliable medical generation.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.