#machine learning
Jun 2026
HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents
HiMPO, a Hindsight-Informed Memory Policy Optimization framework for assigning less-entangled credit to memory-writing actions in long-horizon agents, improves over strong memory-based and RL-based baselines while preserving compressed-context efficiency.
Jiang-Ze Yan, Yili Shen, Wen-Jing Zhang et al.
· arXiv.org · 1 citation