Preprint
Jun 2026
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
A novel policy gradient method is introduced, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them.
Zishang Jiang, Tingyun Li, Jinyi Han et al.
· 0 citations