Skip to content
Review

Medical question answering: A comprehensive multimodal and LLM-driven survey.

Jul 2026 · Computer Methods and Programs in Biomedicine · Vol 285, pp. 109541 · 0 citations · 129 references
Medicine

Abstract

Medical Question Answering (MQA) has emerged as a critical artificial intelligence (AI) capability for supporting clinicians, researchers, and the general public with timely and evidence-based responses to medical queries. Recent advances in natural language processing (NLP), computer vision, and large language models (LLMs) have expanded MQA from text-only systems to multimodal frameworks. This survey aims to provide a comprehensive and structured review of MQA systems, covering both text and image-based approaches. We present a systematic review of MQA literature, including applications, datasets, and modeling paradigms. We introduce a unified taxonomy categorizing MQA systems into scientific, clinical, consumer, and examination-oriented tasks. We also analyze representative datasets for text-based and vision-based question answering, focusing on data sources, annotation strategies, task formulations, and evaluation protocols. Furthermore, we review methodological developments ranging from classical and transformer-based models to multimodal vision-language systems and LLM-driven approaches. The analysis highlights a rapid evolution of MQA systems toward multimodal and LLM-based frameworks, particularly in medical visual question answering. Existing datasets and models demonstrate strong progress but also reveal limitations in generalization, reasoning, and real-world clinical applicability. Key challenges remain, including reliability, hallucination, explainability, fairness, and clinical safety. This survey identifies open research directions such as improved data quality, knowledge-grounded reasoning, trustworthy evaluation, and real-world deployment. The study provides a comprehensive reference and roadmap for developing reliable and clinically applicable MQA systems.

View source

Similar papers

Open access Jul 2026

Enhancing medical Q&A systems with multimodal knowledge graphs and dual-layer attention mechanisms

Medical intelligent question-answering (QA) systems have become important tools for improving the efficiency of healthcare services, and recent research has increasingly emphasized performance optimization and multimodal integration. However, existing systems still face several challenges in intent recognition, entity extraction, and multimodal knowledge fusion, particularly reduced accuracy in multi-label classification, heavy reliance on large-scale annotated data, and limited support for cross-modal retrieval. To address these issues, this study proposes a medical intelligent QA framework that integrates a dual-layer attention mechanism, a large language model, and a multimodal medical knowledge graph to improve system understanding and response generation in complex clinical scenarios. Specifically, we develop a text-based intent recognition model with a dual-layer attention architecture, in which a global contextual attention module is introduced to capture long-range semantic dependencies and improve multi-label classification performance. In addition, an instruction-tuned large language model is employed for zero-shot medical entity recognition, thereby reducing dependence on manually annotated datasets. Building on this foundation, we construct a multimodal medical knowledge graph comprising more than 15,000 associated medical images and develop a visualization-oriented retrieval interface using Flask and ECharts. Experimental results show that the proposed intent recognition model achieves a peak Micro-F1 of 94.42% on multiple benchmark datasets, outperforming several baseline methods. The LLM-based entity recognition module achieved competitive recall in medical entity extraction, demonstrating strong capability in identifying medical entities. User evaluation results further indicate that the system is effective and practical across a variety of medical query types. This study provides a feasible framework for advancing medical QA systems through improved intent recognition, low-resource entity extraction, and multimodal knowledge integration.

Guoqiang Qiu, Qingni Yuan, Yi Wang et al. · 0 citations
Review Open access Jul 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.

Qi Peng, Jiatong Li, Sirui Huang et al. · 4 citations
Open access Aug 2026

Retrieval-augmented generation for medical question answering: a multi-metric performance evaluation

The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.

Yunus Kökver · 0 citations
Open access Jul 2026

The potential of LLMs in generating questions and answers with EHRs

Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.

Yunqi Zhu, Wen Tang, Huayu Yang et al. · 1 citation
Review Open access Jul 2026

Tutorial: guidance on the use of large language models for medical research

This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

Qiao Jin, Nicholas Wan, Robert Leaman et al. · 1 citation