Skip to content

Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach

Jul 2026 · arXiv.org · Vol abs/2607.15661 · 0 citations · 27 references
Computer Science

TL;DR

This work introduces MergeMedBench, a comprehensive benchmark spanning eight imaging modalities and diverse clinical task types, comprising 16 LoRA fine-tuned models built upon two mainstream architectures and proposes winner-take-all, a simple and hyperparameter-free approach that retains only the most dominant parameters across expert models.

Abstract

Large vision-language models (LVLMs) can be adapted to specialized medical imaging tasks via parameter-efficient fine-tuning approaches such as low-rank adaptation (LoRA), leading to a growing ecosystem of expert models tailored to specific imaging modalities and clinical scenarios. However, deploying multiple expert LVLMs in practice incurs substantial computational and operational overhead. Model merging provides a promising solution by consolidating multiple experts into a single model without retraining, yet it remains largely unexplored in the medical domain. In this work, we present the first systematic study of model merging for medical LVLMs. We introduce MergeMedBench, a comprehensive benchmark spanning eight imaging modalities and diverse clinical task types, comprising 16 LoRA fine-tuned models built upon two mainstream architectures. We conduct an extensive evaluation of existing merging methods and further propose winner-take-all, a simple and hyperparameter-free approach that retains only the most dominant parameters across expert models. By preserving the critical parameters that govern model behavior and discarding weaker ones, our method avoids the information dilution inherent in averaging- or alignment-based strategies. Despite its simplicity, winner-take-all consistently outperforms existing approaches, offering both a new perspective on LoRA merging and a strong practical baseline for future research.

View source

Similar papers

Open access Sep 2026

LoRKD: Low-Rank Knowledge Decomposition for Medical Foundation Models.

The widespread adoption of large-scale pre-training techniques has significantly advanced the development of medical foundation models, enabling them to serve as versatile tools across a broad range of medical tasks. However, despite their strong generalization capabilities, medical foundation models pre-trained on lar...

Hao-Lin Li, Yu-Hang Zhou, Zi-Heng Zhao et al. · 0 citations
Preprint Sep 2026

SegBanana: Steering Unified Multimodal Models into Medical Segmenters

SegBanana is proposed, to the authors' knowledge, the first agentic visual generation framework for training-free medical image segmentation and builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation cap...

Xiao-Ye Liang, Ye Yan, Ming-Ze Yin et al. · 0 citations
Preprint Aug 2026

How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

This study finds that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable.

Yi Xiong, L. Gallée, D. Wolf et al. · 1 citation
#large language models Review Open access Sep 2026

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resou...

Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al. · 0 citations
#small language model Preprint Sep 2026

Two-Stage Mixture-of-LoRA for Multi-Task Medical Vision-Language Learning

The proposed framework uses a shared-specific Mixture-of-LoRA architecture comprising one shared LoRA and six task-specific expert LoRAs, together with a two-stage training procedure, and achieves good results in the FLARE 2026 Task 3 test sets.

Zhang-Hao Chen, Yuan-Yuan Li, Zhen-Yu Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.