Skip to content

Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

Sep 2026 · 0 citations · 19 references
Computer Science

TL;DR

SpawnLoRA is evaluated on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and finds that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA, demonstrating that structural separation inside experts provides benefits beyond routing or rank expansion alone.

Abstract

Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.

View source

Similar papers

#small language model Book Open access Sep 2026

CARE-MoE: Correlation-Aware Expert Placement and Semantic Equivalence Routing for MoE LLM Inference on Edge Devices

CARE-MoE is proposed, an efficient MoE LLM inference framework comprising two core components that balances expert placement by jointly modeling co-activation correlation and hot–cold drift, preventing overload from correlated experts and enabling low-cost adaptive rebalancing.

Zhen-Yu Wang, Wei Li, Ao Ren et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing...

Yuan-Yi Wang, Yang-Gan Gu, Su Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, a...

Ahin Lee, Sehyun Yun, Joonha Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs

Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data,...

Zu-Kang Xu, Zhi-Xiong Zhao, Xing Hu et al. · 0 citations
Preprint Aug 2026

MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning

Experimental results show that routing-aware collaboration consistently improves personalized performance compared to conventional federated averaging and local training, while maintaining the same communication cost, and shows that client-centric and expert-centric clustering provides an effective and scalable approac...

Ankita Sharma, B. Farahani, S. Moosavi et al. · 0 citations
Preprint Aug 2026

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuni...

Ali Janati, K. El Maghraoui, Cheng-Ke Zou et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.