Aug 2026· ACM Transactions on Intelligent Systems and Technology· 0 citations· 74 references
TL;DR
ReGAP formulates follow-up question generation as a sequential intervention planning problem, and uses Monte Carlo Tree Search to compare candidate intervention strategies over future dialogue trajectories, and further incorporates experience priors to improve planning efficiency and stability.
Abstract
Conversational question generation (CQG) has shown promise for supporting learners in text-based dialogue. Existing methods mainly optimize question relevance, fluency, and answer consistency, but pay less attention to whether follow-up questions help learners revise incomplete or selective use of evidence. In practice, learners may rely on only a subset of task-relevant evidence, leading to incomplete or overgeneralized interpretations of the source content. To address this problem, we propose Reasoning-Guided Action Planning (ReGAP), a framework for evidence-aligned CQG. ReGAP formulates follow-up question generation as a sequential intervention planning problem. It first estimates the learner's evidence-utilization state from the dialogue history, then constructs structured intervention actions by pairing target evidence with reasoning operations. Based on this state-action representation, ReGAP uses Monte Carlo Tree Search to compare candidate intervention strategies over future dialogue trajectories, and further incorporates experience priors to improve planning efficiency and stability. Experiments on multiple dialogue datasets show that ReGAP consistently improves evidence-aligned reasoning over strong baselines. Expert annotation and real-user studies further demonstrate that ReGAP generates questions that better guide learners toward task-relevant evidence while maintaining conversational fluency and positive user experience. These results suggest that planning-based conversational intervention is a promising direction for evidence-aligned CQG.
The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
The results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.
A rule-based memory framework that induces reusable logical rules from historical interactions to guide both evidence retrieval and reasoning, and constructs natural-language Horn clauses from conversations and validates them via a Rule Perplexity Consistency (RPC) mechanism.
Xing-Yuan Zeng, Zuo-Han Wu, Quanming Yao et al.· 0 citations
GSC-QA (Goal-based Sequential Conversation QA), a framework that integrates three complementary components into a unified enterprise dialogue architecture that combines retrieval, instruction enforcement, and goal persistence in a single coordinated loop built on LangGraph, is introduced.
Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing, and incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies.
Lior Baruch, Moshe Butman, K. Bar et al.· 2 citations
Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrieve and use earlier evidence across dialogue, files, and workspace state. However, agents can generate plausible answers even when access to that histo...
Yu Liu, Wen-Xiao Zhang, Cheng Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.