Skip to content

Retrieval-Guided Fine-Tuning as Noisy Estimation: Risk bounds and Architectural Analysis

Sep 2026 · 0 citations · 26 references
Computer Science Mathematics

TL;DR

Under homoscedastic retrieval noise, it is shown that retrieval failure decays exponentially with task separation relative to noise, and explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training are derived.

Abstract

Retrieval-Guided Fine-Tuning (RAG-FT) incorporates retrieved data directly into the training objective, but the statistical consequences of noisy retrieval during training remain theoretically undercharacterized. We study this question by modeling RAG-FT as an estimation problem in a multi-task linear regression framework, using an OLS proxy for single-layer linear self-attention to obtain finite-sample risk bounds. Under homoscedastic retrieval noise, we show that retrieval failure decays exponentially with task separation relative to noise, and derive explicit finite-sample conditions under which RAG-FT achieves lower risk than both target-only and full-corpus training. We then introduce a Distance-Proportional Noise (DPN) model, in which retrieval quality degrades with rank, and compare two estimators under the same retrieval process: the OLS proxy and the literal, uniform-weight forward pass of linear self-attention. We prove that the attention estimator's bias diverges as $\Theta(n^{2q})$ even under exact retrieval, while OLS risk remains $\Theta(d/n)$ for every noise exponent $q>0$. These results locate the instability not in noisy retrieval itself, but in the fixed, unweighted aggregation of the literal LSA forward pass, which reweighting by reliability empirically removes. We validate the predicted rate separation through direct simulation of the DPN model.

View source

Similar papers

Preprint Aug 2026

The RAT: A Unified Bayesian Model for RAG Evaluation

A Bayesian evaluation framework is introduced that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow, and extends to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited hum...

Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval

This work proposes PAO (Positive-Advantage-Only), a selective RL optimization method that selectively applies gradient updates only to retrieved items with positive advantages, effectively pulling query embed- dings toward high-reward regions while preserving global topo- logical stability.

Shao-Wei Wei, Chong Huang, Songtao Fang et al. · 0 citations
Preprint Aug 2026

Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models

Differentially private zeroth-order optimization (DP-ZO) enables memory-efficient private fine-tuning of large language models using only forward evaluations. Existing aggregation-based DP-ZO methods reconstruct model updates at a fixed scale, ignoring that the strength of useful signals varies throughout training. Con...

Le-Le Zheng, Wei-Feng Kong, Xinyi Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine-Tuning is proposed, a simple method that augments SFT with online, trajectory-level weighting and offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only 2 online rollouts.

Hai-Quan Wen, Yiwei He, Bei Peng et al. · 0 citations

STAR: Structure-Aware Adaptive Retrieval for RAG

STAR is presented, a structure-aware adaptive retrieval framework for RAG that treats this mismatch as a problem of diagnosing evidence sufficiency and benefits from a control signal that preserves structurally distinct insufficiency patterns rather than collapsing them into a single scalar confidence estimate.

Yeowon Jeon, Chong-kwon Kim, Y. Choi · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.