Skip to content

Author

Adnan Khan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Towards Multimodal Retrieval-Augmented Generation for Medical Visual Question Answering

A novel multimodal RAG framework tailored for MedVQA is proposed, which leverages multimodal data, including medical images, reports, and generated captions, to provide more accurate clinical answers, and introduces a training paradigm that uses captions as auxiliary supervision, enhancing cross-modal alignment via contrastive learning.

Mai A. Shaaban, M. Zarei, Adnan Khan et al. · 0 citations