Skip to content

Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning

Sep 2026 · 1 citation · 33 references
Computer Science

TL;DR

The analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value, and show that stricter feasibility need not improve final task performance.

Abstract

Orthogonality in a LoRA factor does not by itself specify what the composed update protects: the answer depends on the task-start state, the parameterization, and the realized optimizer displacement. We formalize this question through Bilinear Optimization Divergence (BOD), an anchor-relative diagnostic of effective-update response on selected historical features. The finite-step analysis distinguishes two cases. In a shared adapter, protecting the routing displacement leaves a learned-anchor residual through the changing companion factor. In a fresh zero-output block, a feasible routing state can protect the composed update while both current factors remain trainable. These conditions yield Semi-Frozen Orthogonal Routing (SFOR) for shared adapters and current-block hard protection for cumulative O-LoRA; Weight Residual Projection (WRP) enforces the required displacement after the optimizer step. Controlled two-task traces verify the predicted residual paths, reducing normalized historical response from 19.12% to 0.005% in the shared family and from 7.72% to 0.002% in the cumulative family. Four-task experiments on Qwen3-8B characterize the resulting trade-offs: SFOR improves backward transfer (BWT) from -2.47 to -0.86 with nearly unchanged average accuracy (AA), while O-LoRA hard protection improves three-order mean AA from 80.27% to 81.30% and forgetting measure (FM) from 2.20 to 0.43. Component controls also show that stricter feasibility need not improve final task performance. Together, the analysis and evidence provide an architecture-conditioned account of which constraint to enforce, how to enforce it, and how to interpret its empirical value.

View source

Similar papers

#machine learning Preprint Sep 2026

Jacobian Rank Collapse in Decision-Focused Learning

Decision-focused learning (DFL) trains predictors through downstream objectives, but a different loss need not provide an independent parameter-update direction. We characterize this restriction through the predictor Jacobian, using sparse index tracking to distinguish the covariance entries read by the optimizer from...

Ao-Jie Yuan, Hai-Yu Zhang, Zi-Jian Su · 0 citations
#machine learning Preprint Sep 2026

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

FCCA, which estimates the signed input--error cross-covariance, whitens it with diagonal Fisher moments, truncates it in the resulting local metric, maps the selected directions back, and applies thin QR to obtain stable core coordinates, shows that a carefully selected fixed span can recover most of the benefit of mov...

Wen-Song Ye, Zhan-Ming Shen, Zhiqing Xiao et al. · 1 citation · ⚡1
#machine learning Preprint Sep 2026

Reducing the Adaptation Gap Through Reachable Fisher Geometry

Parameter-efficient fine-tuning (PEFT) determines not only how many parameters are trained, but also which local directions a model can move in, so similar adapters can affect subgroup losses differently. Since curvature matrices are infeasible to form at adapter scale, scalar summaries such as the Fisher trace are oft...

Wasif Jalal, Sachin Deb, Asif Salekin · 0 citations
#machine learning Preprint Oct 2026

MuLoRA: Spectrally Balanced Low-Rank Adaptation for Continual Learning

Low-rank adaptation (LoRA) provides a parameter-efficient approach to continual learning, but its nominal rank can conceal a loss of effective adaptation capacity. We identify \emph{spectral plasticity collapse}: during sequential adaptation, update energy becomes concentrated in a small subset of singular modes, leavi...

Jun-Kang Liu · 0 citations
Open access Sep 2026

DWAT: Density-Weighted Adversarial Training for Robustness Beyond the Training Perturbation Budget

Deep neural networks (DNNs) are widely deployed in safety-critical applications such as medical diagnosis and autonomous driving. Adversarial training (AT) is among the most effective defenses, casting robust optimization as a min–max problem over a defender-specified ℓp-ball of fixed radius ϵ. Bounded defenses of this...

Jie-Ying Huang, Rui-Ming Zhu, Jia Xu et al. · 0 citations
Preprint Aug 2026

CG4AI: A Column Generation Framework for Training AI Models Under Constraints

This work proposes CG4AI, a framework that builds a convex combination of AI models while enforcing linear constraints on the combined output, and applies it to digit classification on MNIST and the multi-commodity flow problem, where link capacity constraints are enforced on neural-network routing predictors.

Youcef Magnouche, Abderrahmane Driouch, Sébastien Martin et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.