Skip to content

Capacity and Redundancy Trade-offs in Multi-Task Learning

Jul 2026 · arXiv.org · Vol abs/2607.16554 · 0 citations · 48 references
Computer Science Mathematics

TL;DR

A Capacity--Redundancy (CR) identity is investigated that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation.

Abstract

In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling $\Delta$ from validation residual correlations, showing that clustered LoRA substantially reduces $\widehat{\Delta}$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.

View source

Similar papers

Preprint Aug 2026

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

Model merging provides an efficient way to construct multi-task generalist models without additional training, but its performance often degrades under severe task interference. Task interference in model merging primarily stems from \textit{superposition}, where task-specific features become entangled within the param...

Yi-Hang Zhang, Shengen Sun, Junbo Wen et al. · 0 citations
Preprint Aug 2026

Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction

MemMTL, a multi-task dense prediction framework that estimates a compact task state from global visual context and refines it through a learnable task-state prototype memory, is proposed.

Yang-Yang Xu, Haobo Yuan, Yu-Zhu Wang et al. · 0 citations
Preprint Aug 2026

CD-LoRA: Consistency-Driven Low-Rank Adaptation for Multi-Task Fine-Tuning

By eliminating routers entirely, CD-LoRA employs a consistency-driven alignment mechanism to enforce representation congruence across tasks in a shared low-rank space, which fosters robust, task-agnostic features without explicit partitioning overhead.

Qian Zha, Jinda Liu, Yuan Wu et al. · 1 citation
Preprint Aug 2026

SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexi...

Kejian Zhu, Zhuo-Ran Jin, Shangqing Tu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.