MedEvidence-RAG for Joint Evidence Retrieval and Generator Alignment in Medical Visual Question Answering and Report Generation
Abstract
Medical vision-language models remain vulnerable to unsupported findings because their outputs are weakly grounded in images and external clinical evidence. We propose MedEvidence-RAG, which couples a mixture-of-experts (MoE) multimodal retriever with retrieval-conditioned generator alignment. The retriever uses a SigLIP-style objective to match multimodal image-question queries with reports, while supervised fine-tuning (SFT) and direct preference optimization (DPO) teach the generator to use the same retrieved context available at inference. Across IU-Xray, MIMIC-CXR, and Quilt-1M, MedEvidence-RAG improves medical visual question answering (VQA), reaching 89.98 F1 and 91.29 area under the receiver operating characteristic curve (AUC) on MIMIC-CXR. It also achieves the best ROUGE-L on both report-generation benchmarks. Ablations show that retrieval and alignment provide complementary gains, and retrieval analysis confirms more relevant evidence than a Contrastive Language-Image Pre-training (CLIP)-based retriever. These results support joint optimization of evidence acquisition and evidence use for reliable medical generation.