Skip to content

Rethinking Expert Training for Model Merging with Prompt Learning

Jul 2026 · arXiv.org · Vol abs/2607.24465 · 0 citations · 47 references
Computer Science

TL;DR

Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder, is introduced, a two-stage training strategy that reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility.

Abstract

Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work, we revisit expert training for model merging. We first show that prompt-based adaptation provides a strong baseline: independently learned prompts can be exploited across tasks while keeping the backbone fixed, avoiding the interference introduced by weight merging. Building on this observation, we introduce Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder. This reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility. Experiments across multiple CLIP architectures, full fine-tuning, and LoRA experts show that DTEs consistently improve merged performance of standard merging approaches and remain effective even when combining heterogeneous sets of experts.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning

Multi-Objective In-context Knowledge Editing (MO-IKE), a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process, enabling more balanced and globally coherent prompt construction.

Xu-Zhong Wang, Maiqi Jiang, Tejal Nair et al. · 1 citation
Preprint Jul 2026

Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning

Evidence is provided that cross-scale heterogeneous fusion can succeed without explicit semantic alignment when the donor contribution is sufficiently concentrated and carefully selected, and that activation-guided extraction improves the quality of the transferable donor slice while preserving the small-ratio fusion r...

Jiahe Fan, Si Chen, Yinghao Hou et al. · 0 citations
Aug 2026

DR-EFT: Exploring and reloading domain-representative experts for the memory-constrained fine-tuning of MoE large models.

An algorithm framework named DR-EFT (Domain-Representative Experts for Fine-Tuning), which explores and loads the domain-representative experts for subsequent retraining and reincorporation and demonstrates robustness through validations on popular MoE LLMs, including Qwen, DeepSeek, and Ernie.

Zhaomeng Cheng, Zhong Ji, Yan Zhang et al. · 0 citations
#natural language process... Preprint Aug 2026

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms

This work compares three fusion paradigms by the artefacts they reuse and suggests that Merge when experts already exist and cheap fusion is paramount; Mix RL when training a unified model without experts, with domain proportions adjusted for cross-domain transfer; and MOPD when preserving domain-specific gains matters...

Sicheng Wu, Kai Yang, Yuchen Cai et al. · 0 citations
Conference Open access Sep 2026

MOBO: A Merging-Oriented Bi-Level Optimization Framework for Class Incremental Learning

Class-Incremental Learning (CIL) aims to enable models to sequentially learn new tasks while retaining knowledge from previous ones. Recently, merging-based pre-trained CIL methods have gained significant attention due to their competitive performance and high inference efficiency. However, most existing approaches dec...

Si-Yu Zhang, Wen Wang, Wen-Ju Sun et al. · 0 citations
#natural language process... Preprint Sep 2026

Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist

Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometri...

Ming Cheng, Jiaying Gong, Hoda Eldardiry · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.