Jul 2026· ACM Transactions on Computing for Healthcare· 1 citation· 54 references
TL;DR
This work introduces IndicMedQA, a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues, and creates a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach.
Abstract
In the rapidly advancing field of AI-driven telehealth services, effective medical communication in India remains a significant challenge due to the language barrier, as most of the population is not proficient in English. Additionally, many patients struggle to accurately describe their medical conditions using text alone. As a result, an essential feature of any telehealth service is the ability to supplement textual queries with medical images, enabling doctors to conduct a more careful analysis and provide well-informed diagnoses and treatment recommendations. In this work, we introduce IndicMedQA , a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues. To support this, we create a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach. This dataset facilitates medical understanding for every input query and its associated medical image—the output is a detailed patient summary, symptom analysis, probable conditions, additional findings, and severity assessment with medical precision. Our framework significantly enhances personalized healthcare experiences, ensuring context-aware multimodal understanding of patient needs in Indic languages. Extensive experiments demonstrate that IndicMedQA surpasses all baselines, establishing a new benchmark for Indic AI in healthcare. Disclaimer: This work contains medical images that depict the subject matter of the study, which may be disturbing to some readers.
A novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria reveals several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts.
Tobi Olatunji, C. Aka, C. Okocha et al.· medRxiv· 0 citations
Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on English, limiting its relevance to linguistically diverse patients and clinicians. Recent multilingual medical VQA benchmarks show that large vision-language models (LVLMs) degrade in non-English languages, but lack a fine-grained analysis of how cross-lingual variation affects the distinct capabilities that medical VQA requires. To this end, we construct a multilingual medical VQA benchmark over eight languages, organized into four representative scenarios that isolate the core capabilities medical VQA requires. Evaluating five open- and closed-source LVLMs, we find that cross-lingual degradation is not uniform but highly scenario-dependent. We therefore propose MedVL-XLRepE, a training-free scenario-aware representation engineering method, leveraging LVLMs'superior English medical VQA capability to steer non-English representations toward their English counterparts at inference time. Across three LVLMs and eight languages, MedVL-XLRepE consistently mitigates cross-lingual degradation, with gains of up to 6.33\%.
Jingbo Wang, Sendong Zhao, Haochun Wang et al.· 0 citations
The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.
Lai Wei, Yuchao Chen, Zhenbiao Cao et al.· 0 citations
A large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital, MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation.
Runhan Shi, Quan Zhou, Yuqian Xu et al.· 0 citations
Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer grade hardware, and Gaokerena-R, a novel family of compact Persian medical language models optimized for deployment on consumer grade hardware, are presented.
The research came up with a bilingual natural language processing (NLP) model of a patient-focused healthcare system that is geared towards improving healthcare communication in low-resource settings, based on both the English and Yoruba languages. The system combines multilingual transformer-based text processing, speech-to-text and text-to-speech modules, retrieval-augmented response generation, and offline-compatible deployment to provide culturally adaptive and personalized health information. Assessment was made based on conventional performance measures such as accuracy, cultural relevance, readability, and response quality through pilot testing and comparison against other existing bilingual models of healthcare delivery. The experimental findings indicated a general classification accuracy of 91%, which revealed a high language interpretation and response-generating ability. The system also achieved cultural relevance of 40% by adapting to indigenous healthcare communication requirements and readability of 91% by simplifying complex medical information into patient-friendly explanations. Analysis indicated that linguistic inclusiveness and accessibility were enhanced as compared to the current monolingual practices. The results confirm that bilingual NLP models have the potential to significantly reduce communication barriers, enhance patient comprehension, and promote equitable healthcare practices in bilingual and underserved communities. Comparative analysis confirmed that the proposed bilingual NLP model outperforms existing monolingual and rule-based systems in linguistic inclusiveness and accessibility. The results validate that bilingual NLP models can significantly reduce communication barriers, enhance patient comprehension, and promote equitable healthcare delivery in multilingual, underserved communities.