Skip to content

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Jul 2026 · arXiv.org · Vol abs/2607.11172 · 1 citation · 31 references
Computer Science

TL;DR

STAMP is proposed, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it.

Abstract

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

View source

Similar papers

Preprint Aug 2026

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Answer-Backtracked Credit Assignment (ABC) is proposed, a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant a...

Yi-Jun Lu, Rui Ye, Jia-Jun Wang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored o...

Shubham Gandhi, Saurabh Goyal, K. Kate et al. · 1 citation
Preprint Aug 2026

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Evidence Anchors are constructed, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path, and SSPO, which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exc...

Haoze Wu, Chu-Qiao Kuang, Tian-Yi Zhuang et al. · 1 citation
#artificial intelligence Preprint Aug 2026

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

VICT (VerifierInstrumented Credit Tracing), a training-time interface that exposes executable or evidence backed atoms and traces them back to actions through dependency-valid proof edges, improves substantially over outcome-only training and achieves strong performance alongside recent fine-grained credit methods.

Peng-Cheng Li, Zhengyang Zhang, Dong-Xu Zhang et al. · 0 citations
#natural language process... Preprint Sep 2026

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by...

Qiang Zhang, Rui-Xue Ding, Fanrui Zhang et al. · 1 citation
Preprint Aug 2026

TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents

Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult. Existing approaches either rely on process evaluators, which incur annotation and inference costs, or derive step-level credit from succes...

Huan-Xi Zhang, Ming-Ju Chen, Dongxu Zhou et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.