Skip to content
Conference

MedQwen-VGR1: Multimodal Vision-Guided Reasoning Model for Med-VQA, Temporal Diagnosis and Drug Interaction Analysis in Medical Imaging and Pathology

Jul 2026 · 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET) · pp. 1-6 · 0 citations · 16 references

Abstract

MedQwen-VGR1 is a novel multimodal visionlanguage model (VLM) trained for medical visual question answering (Med-VQA), longitudinal temporal diagnosis, and drug interaction analysis across radiology and pathology imagery. The model undergoes a multi-stage training pipeline comprising: (1) Continuous Domain-Adaptive Pretraining on heterogeneous medical image-text corpora; (2) Supervised FineTuning on expert-annotated multimodal datasets spanning sequential imaging conversations; (3) Human Preference Alignment via Direct Preference Optimization (DPO) and Group Relative Preference Optimization (GRPO) applied to Vision-Guided Chain-of-Thought (CoT) Reasoning trajectories; and (4) Smart Memory Module integration enables Cross-Visit Context Retention and Comparative Reasoning on sequential or temporal pathologies. MedQwen-VGR1 processes sequential medical images and textual queries to generate stepwise, evidence-grounded rationales correlating visual features across timepoints, predict adverse drug interactions from images/metadata, and output calibrated reliability scores. Evaluations on PathVQA, SLAKE, and VQA-RAD demonstrate superior performance over LLaVA-Med (+14.7% accuracy) and BioViL-T (+12.3% temporal reasoning), achieving 82.4% VQA accuracy and 78.6% change detection precision, enabling multi-turn clinical dialogues and automated clinical reports. The framework enables dynamic, multi-turn clinical dialogues with automated medical report synthesis, addressing critical gaps in temporal reasoning, explainability, and clinical safety alignment for realworld AI and safety enhanced clinical decision support systems.

View source

Similar papers

#artificial intelligence Book Open access Jul 2026

MedTVL: Harnessing Vision and Language for Medical Time Series Classification

Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, MedTVL is introduced, a text-guided dual-pathway architecture tailored for MedTS classification that synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences a...

Jiexia Ye, Jia Li, F. Tsung · 0 citations
Sep 2026

Lingshu: Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.

Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in me...

Wei-Wen Xu, H. Chan, Long Li et al. · 0 citations
Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
Preprint Aug 2026

MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA

A lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment is proposed, highlighting its efficiency and potential for clinical deployment.

Hao-Wen Gu, Gen-Sheng Pei, Zeren Sun et al. · 2 citations
Conference Open access Sep 2026

MedVCoT: Bridging the Modality Gap in Medical VQA Through Latent Visual Reasoning

This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-...

Bo Xu, Quan-Hao Zhu, Bo-Lin Zhu et al. · 1 citation
#natural language process... Preprint Sep 2026

MedProb: Probing Internal Representations of Vision-Language Models for Medical Question Answering

Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text g...

E. Nourbakhsh, Ke Yang, Anthony Rios · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.