Jul 2026· 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)· pp. 1-6· 0 citations· 16 references
Abstract
MedQwen-VGR1 is a novel multimodal visionlanguage model (VLM) trained for medical visual question answering (Med-VQA), longitudinal temporal diagnosis, and drug interaction analysis across radiology and pathology imagery. The model undergoes a multi-stage training pipeline comprising: (1) Continuous Domain-Adaptive Pretraining on heterogeneous medical image-text corpora; (2) Supervised FineTuning on expert-annotated multimodal datasets spanning sequential imaging conversations; (3) Human Preference Alignment via Direct Preference Optimization (DPO) and Group Relative Preference Optimization (GRPO) applied to Vision-Guided Chain-of-Thought (CoT) Reasoning trajectories; and (4) Smart Memory Module integration enables Cross-Visit Context Retention and Comparative Reasoning on sequential or temporal pathologies. MedQwen-VGR1 processes sequential medical images and textual queries to generate stepwise, evidence-grounded rationales correlating visual features across timepoints, predict adverse drug interactions from images/metadata, and output calibrated reliability scores. Evaluations on PathVQA, SLAKE, and VQA-RAD demonstrate superior performance over LLaVA-Med (+14.7% accuracy) and BioViL-T (+12.3% temporal reasoning), achieving 82.4% VQA accuracy and 78.6% change detection precision, enabling multi-turn clinical dialogues and automated clinical reports. The framework enables dynamic, multi-turn clinical dialogues with automated medical report synthesis, addressing critical gaps in temporal reasoning, explainability, and clinical safety alignment for realworld AI and safety enhanced clinical decision support systems.
Inspired by diagnostic practice synergizing numerical assessment, visual inspection and clinical context, MedTVL is introduced, a text-guided dual-pathway architecture tailored for MedTS classification that synergizes a convolution-based temporal pathway for fine-grained temporal dynamics from raw numerical sequences a...
Jiexia Ye, Jia Li, F. Tsung· Proceedings of the 32nd ACM...· 0 citations
Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in me...
Wei-Wen Xu, H. Chan, Long Li et al.· IEEE Transactions on Pattern...· 0 citations
This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.
Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al.· 0 citations
A lightweight framework that leverages a memory bank to augment DCT-based low-frequency features and employs graph-enhanced cross-attention for effective visual-textual alignment is proposed, highlighting its efficiency and potential for clinical deployment.
Hao-Wen Gu, Gen-Sheng Pei, Zeren Sun et al.· 2 citations
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-...
Bo Xu, Quan-Hao Zhu, Bo-Lin Zhu et al.· Proceedings of the Thirty-Fi...· 1 citation
Medical visual question answering (Med-VQA) is often assumed to require medical fine-tuning, large models, or complex multi-agent pipelines. We revisit this assumption with \textbf{MedProb}, a lightweight probing framework that predicts multiple-choice Med-VQA answers from frozen VLM representations without free-text g...
E. Nourbakhsh, Ke Yang, Anthony Rios· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.