Skip to content

Author

Haowei Liu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning

HindSearch is introduced, a hindsight self-distillation procedure for GRPO: after each rollout, a frozen judge writes a short critique of every failed trajectory using the gold answer, and the critique supplies an auxiliary on-policy distillation signal on the student's search actions.

Hao-Wei Liu, Jiamian Wang, Hsin-Tai Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.