Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing, and incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies.
Abstract
Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead. Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets. By combining this method with Direct Preference Optimization (DPO), we aim to enhance the agent's decision-making capabilities over iterative training cycles. The proposed framework addresses data scarcity and advances the development of more nuanced and effective dialogue systems in goal-oriented domains. Experimental evaluations demonstrate that the PTO framework enhances dialogue agents'performance in goal-oriented conversations within the domain of Motivational Interviewing (MI). Models trained with PTO consistently outperformed the baseline in key metrics such as session satisfaction and working alliance. Additionally, incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies, with deeper look-ahead configurations yielding the most stable and high-scoring results.
ReGAP formulates follow-up question generation as a sequential intervention planning problem, and uses Monte Carlo Tree Search to compare candidate intervention strategies over future dialogue trajectories, and further incorporates experience priors to improve planning efficiency and stability.
Wanqiang Wang, Long-Zhu He, Peng-Peng Zhou et al.· ACM Transactions on Intellig...· 0 citations
DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT), suggests that combining stage-structured dialogue with learned strategy selection is a promising approach for AI cou...
Qi Zhang, Heajun An, P. Dumaru et al.· 0 citations
The results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.
Information analysis recommendation differs from conversational recommender systems (CRS) because relevance changes with the decision phase. The same event may support observation, interpretation, option selection, or action feedback, yet most large language model (LLM)-agent CRS represent dialogue state as intent and...
Chao-Yang Li, Yiwei Lu, Bo Huang et al.· Electronics· 0 citations
This work introduces StageWell, a process-aligned Chinese corpus for positive psychology dialogue together with HQS, a structured protocol for data construction and evaluation, and highlights the value of modeling supportive dialogue as a structured multi-turn support process rather than as single-turn response generat...
Yuxun Wang, Zihan Lin, Bo Wang et al.· 0 citations