Skip to content

Learning to Coach for Experiential Learning

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

This work proposes Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory, and studies two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances.

Abstract

Language models can learn from experience, but raw solution trajectories are often too long and noisy to provide effective guidance. In this work, we propose Learning to Coach (L2C), a framework that trains a dedicated LLM-as-a-Coach to extract actionable experiential knowledge from an actor model's previous trajectory. The actor remains frozen, while the LLM-as-a-Coach is trained to maximize a reward given by the correctness of the actor's guided response. We study two such rewards: a same-instance reward, which improves subsequent responses on the original problem, and a cross-instance reward, which elicits knowledge that transfers to other instances. Across mathematical reasoning and interactive text-games, L2C consistently outperforms self-refinement and an untrained LLM-as-a-Coach. Running experiential learning for more iterations further improves accuracy and uses additional inference compute more effectively than enlarging the actor's decoding budget. The trained LLM-as-a-Coach also transfers to out-of-distribution tasks and adapts its guidance to the specific actor it coaches.

View source

Similar papers

#machine learning Preprint Aug 2026

Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL

This work extends contrastive reinforcement learning (CRL), a prototypical self-supervised method, to operate over action chunks, and finds that this results in large, pervasive gains across established offline and online benchmarks: +31.7% and +93.1% across 18 and 11 environments respectively.

Michal Korniak, Kamil Dybek, Benjamin Eysenbach et al. · 1 citation · ⚡1
Preprint Aug 2026

Chain-of-Experience for Continual LLM Improvement

This study studies how LLMs learn from iterative experience at test time, a setting the authors refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference.

Hao-Qin Tu, Yun-Hao Fang, Yizhong Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PoEM: Predicting RL Outcomes from Existing Policies

PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards, is introduced by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards.

K. Hamidieh, G. Daras, Antonio Torralba · 0 citations
#artificial intelligence Preprint Sep 2026

RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and wh...

Michael Kirchhof, Eleonora Gualdoni, Andrew Szot et al. · 0 citations
#machine learning Preprint Sep 2026

Learning to Optimize through Solver-Grounded Self-Play

Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al. · 0 citations
#machine learning Preprint Sep 2026

Cliff: Learning Process Rewards from the First Mistake

Cliff, a reward shaping strategy that utilizes an off-the-shelf LLM as a teacher to identify the first mistake in each rollout, is proposed and established as a simple, general and effective approach for improving RLVR with richer, fine-grained supervision.

Pei-Xuan Han, Runnan Wang, Ketan Ramaneti et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.