Next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents and improves the nine-domain $\tau$-family average across three independently trained runs.
Abstract
User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repair. The next user turn is more than context: it also provides noisy, temporally local evidence about the preceding user-to-user segment. We introduce \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}), which aligns each reaction with that segment, derives a locally normalized reaction advantage, and adds it to verified terminal outcome advantage without an extra critic or rollout. Against an outcome-only Interactive GRPO control matched in simulator, visible dialogue, initialization, rollout, and optimization, \textsc{FACA} improves the nine-domain $\tau$-family average across three independently trained runs by 5.91 and 10.22 percentage points at 8B and 14B, respectively. Gains concentrate in Telecom; at 8B, randomizing reaction polarity removes the Telecom gain. The same ordering holds zero-shot on Pare-Bench and Co-Gym. These results demonstrate that next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents.
Proactive dialogue requires agents to continually adapt their policies to user feedback while progressing toward task objectives over multiple turns. To move beyond imitation learning on static datasets, recent approaches use user simulators to collect interactive data for policy optimization. However, many simulators...
Ming-Hui Ma, Meng-Qi Chen, Bin Guo et al.· 0 citations
This work introduces UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive, and establishes per-turn intent control as a complementary dimension to response fidelity in user simulation.
Bo Wang, Ruixing Zhang, Yunqi Liu et al.· 0 citations
We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a modified system, because a different response changes every...
Kasra Hosseini, Wen-Sen Cheng, Marco-Andrea Buchmann et al.· 0 citations
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limite...
Shuang Sun, Guo-Xin Chen, Fan-Zhen Meng et al.· 0 citations
The findings suggest that the primary benefit of human--agent interfaces may be reducing interaction effort rather than improving speed, and that delegation reflects who the user is more than what the task demands.
Gavin Raine Dizon, Tyrone Justin Sta Maria, Jordan Aiko Deja et al.· 0 citations
It is suggested that imitating full trajectories helps with playability, while turn-level and teacher-guided training usually improve decision-making and increase the overall score, and small models are performant simply by using careful curation strategies rather than aggressive changes.
Syed Mahbubul Huq, P. Madhyastha· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.