Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting, and further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustne...
Rui-Ke Cao, Fu-Gen Yao, Liang Dong et al.· 0 citations
Long-running LLM agents require memory that persists and evolves across sessions. Text-based memory retrieves and reconstructs past interactions at every query, making long-horizon performance increasingly dependent on retrieval quality and contextual reasoning as histories grow. Parametric memory encodes experience di...
Fan-Yu Zhao, Rui-Ke Cao, Liang Dong et al.· 0 citations
MemCalib-RL is proposed, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation and achieves the best overall performance while better balancing over-use and under-use.
Rui-Ke Cao, Fan-Yu Zhao, Fu-Gen Yao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.