This work proposes Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory, and studies two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances.
Abstract
Language models can learn from experience, but raw solution trajectories are often too long and noisy to provide effective guidance. In this work, we propose Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory. The actor remains frozen, while the LLM-as-a-Coach is trained to maximize a reward given by the correctness of the actor's guided response. We study two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances. Across mathematical reasoning and interactive text-games, L2C consistently outperforms self-refinement and an untrained LLM-as-a-Coach. Running experiential learning for more iterations further improves accuracy and uses additional inference compute more effectively than enlarging the actor's decoding budget. The trained LLM-as-a-Coach also transfers to out-of-distribution tasks and adapts its guidance to the specific actor it coaches.
This work extends contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and finds that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively.
Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al.· 1 citation· ⚡1
This study studies how LLMs learn from iterative experience at test time, a setting the authors refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference.
Hao-Qin Tu, Yun-Hao Fang, Yizhong Wang et al.· 0 citations
PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards, is introduced by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards.
K. Hamidieh, G. Daras, Antonio Torralba· 0 citations
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and wh...
Michael Kirchhof, Eleonora Gualdoni, Andrew Szot et al.· 0 citations
Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al.· 0 citations
Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.
Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al.· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.