Skip to content
Preprint

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Aug 2026 · 0 citations · 50 references
Computer Science

TL;DR

A Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor is proposed.

Abstract

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success. Its effectiveness hinges on the guidance depth: how much of the trajectory to keep. Existing methods treat this depth as a deterministic scalar. Scheduled approaches share one value across samples and ignore per-task heterogeneity; per-sample probing estimates it separately at the cost of extra rollouts. We find that useful guidance occupies a band of depths whose informativeness profile is approximately Gaussian around the band center, rather than concentrating at a single optimal point. We propose Agent-G$^2$, a Gaussian guidance framework that draws the depth per task from a Gaussian whose center and spread are estimated online from rollouts already collected for policy optimization, requiring no probe rollouts or learned depth predictor. The center combines a global baseline with per-cluster difficulty, and the spread tracks within-cluster variance. We evaluate Agent-G$^2$ on ALFWorld and WebShop on Qwen2.5-1.5B / 7B-Instruct. Agent-G$^2$ outperforms the strongest hint-based, hint-free, and Aux-RL baselines on ALFWorld by 2.3 / 3.9 / 7.4 points at under one-third the rollout cost of per-sample probing.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deploy...

Guo-Wei Zou, Hai-Tao Wang, Guo-Xin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning

Empirically, EPIG reduces gradient MSE in cloned-state control, winning in all nine dense continuous-control environments of a 13-environment sweep and recovering the reference gradient direction near-perfectly, and it improves frozen-LLM gradient calibration relative to entropy branching.

Nikita Khomich, L. Hermansson, Ido Hakimi · 0 citations
Preprint Sep 2026

Is One Step Enough for Offline Policy Improvement?

This work studies how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy, and identifies improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and add...

Soohyun Choi, Seonvin Cho, Songnam Hong · 0 citations
#machine learning Preprint Sep 2026

When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

The cross-signal NTK is introduced, a token-level statistic that measures the alignment between reward and distillation gradients at position n and an empirical threshold beyond which naive mixing can lead to persistent training collapse is revealed.

Xin-Ke Jiang, Tao Feng, Zhi-Bang Yang et al. · 0 citations
Preprint Aug 2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al. · 7 citations · ⚡1
Preprint Aug 2026

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Evidence Anchors are constructed, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path, and SSPO, which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exc...

Haoze Wu, Chu-Qiao Kuang, Tian-Yi Zhuang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.