2026· SIGDIAL Conferences· pp. 151-163· 0 citations· 66 references
Computer Science
TL;DR
This paper surveys how conversational grounding is evaluated in task-oriented dialogue in the current era of LLMs and focuses on how conversational grounding is modelled explicitly—using dialogue acts and by modelling the participant mental state.
Abstract
Conversational grounding is the collaborative process through which speakers establish and maintain mutual understanding. It is essential for the success of a dialogue. While it is inherent in human conversations, it remains a challenge for instruction-following Large Language Models (LLM). This paper surveys how conversational grounding is evaluated in task-oriented dialogue in the current era of LLMs. First, we focus on how conversational grounding is modelled explicitly—using dialogue acts and by modelling the participant mental state. Then, we review collaborative tasks that enable the evaluation of conversational grounding implicitly at the global level based on outcomes. Finally, we highlight the limitations of evaluation, notably the current methodology and metrics used, and outline research directions in conversational grounding and its evaluation.
Conversational style in this setting is best understood as a pragmatic orientation toward the user rather than toward the system’s own content, and that the system must be mixed-initiative to manifest a conversational style.
Vrindavan Harrison, M. Walker· Proceedings of the 26th ACM...· 0 citations
This work proposes a pipeline that leverage LLMs as safety detector, editor and evaluator to mitigate undesired behaviour in human-computer dialogues and shows reduction in the unsafe dialogues after revision.
T. Ajayi, M. Arcan, P. Buitelaar· WOCHAT2026: Workshop on Chat...· 0 citations
Immersive and believable NPC dialogue requires characters that feel intentional. They remember specific information about themselves, follow through on their goals, and stay true to their personalities across long conversations. We introduce Post-Thinking, a technique that maintains a rolling reflection trace across ch...
Keegan Carey, Hexi Wang· International Conference on...· 0 citations
The results show that CoRG remains challenging for current agents, even the best agent reaches only 67.0% success rate, leaving one third of references unresolved, and position CoRG as a concrete benchmark for studying how agents search, inspect, and verify information in realistic multi-tool environments.
This study directly test whether that omitted within-conversation context changes answers in a conversation and concerns preceding turns in the same conversation and does not test persistent memory across separate conversations.
This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.
Pei-Rong Li· Scientific Journal of Intell...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.