Skip to content

Author

Jiansheng Wei

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jun 2026

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training

A novel policy gradient method is introduced, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them.

Zishang Jiang, Tingyun Li, Jinyi Han et al. · 0 citations