Skip to content
Open access

Persona-centric Metamorphic Relation Guided Robustness Evaluation for Multi-turn Dialogue Modeling

Aug 2026 · Cognitive Computation · Vol 18 · 0 citations · 64 references

TL;DR

This work discovers persona-centric metamorphic relations to infer test samples from annotated data, without additional annotation cost, and evaluates the robustness of personalized dialogue models regarding persona consistency, revealing that prompt learning is more robust than training from scratch and fine-tuning.

Abstract

Retrieval-based dialogue systems aim to select a proper response according to multi-turn conversational history. Persona-based conversation utilizes prior knowledge to maintain persona consistency, enhancing retrieval accuracy. However, reference-based evaluation relies on high-quality data annotation, which is costly and time-consuming. To address this, we discover persona-centric metamorphic relations to infer test samples from annotated data, without additional annotation cost. Benefiting from this, this work efficiently evaluates the robustness of personalized dialogue models regarding persona consistency. Specifically, we discover three types of metamorphic relations from three aspects: self-persona, partner-persona, and response, to automatically derive new test samples . Then the inherent inference relations between originals and derivatives allow for robustness evaluation. Using this evaluation methodology, our work assesses three widely used training paradigms: non-pretraining, fine-tuning after pre-training, and prompt learning, in personalized dialogue retrieval to observe whether these paradigms are more robust or exhibit the same flaws as the other two paradigms. Our experimental results, based on the three discovered metamorphic relations with consistent outputs reveal that prompt learning is more robust than training from scratch and fine-tuning. While traditional reference-based validation and natural language processing methods achieve competitively high retrieval accuracy (Hits@1 up to 87.4%), the persona consistency of dialogue retrieval systems is just 20.98% when persona descriptions are perturbed using various metamorphic relation-based transformations.

Read PDF

Similar papers

Jul 2026

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

This work analyzes real chatbot failures to identify six recurring mechanisms and defines six controlled evaluation modes in Hy-MultiTurn, a Chinese benchmark for deep multi-turn dialogue understanding, which shows that Hy-MultiTurn is broadly challenging.

Eileen Ye, Ji-Hua Tao, Yao-Ming Li et al. · 0 citations
#natural language process... Preprint Aug 2026

Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems

An unsupervised fine-tuning pipeline that harvests reasoning trajectories via in-context learning inference via in-context learning inference is proposed, enabling Large Language Models (LLMs) to access external knowledge and produce factual responses.

Mark A. Ferro, Oier López de Lacalle · 0 citations
#large language models Book Open access Sep 2026

Overview and Analysis of the RecSys Challenge 2026: Conversational Music Recommendation

The RecSys Challenge 2026 studies conversational music recommendation as a joint item recommendation and response generation problem: given a multi-turn dialogue, systems must retrieve relevant tracks from a large catalog and produce a grounded natural-language response. This paper presents the challenge task, dataset,...

Seungheon Doh, Sergio Oramas, B. Sguerra et al. · 0 citations
Open access Sep 2026

Boundary-Centered Explainability for Dialogue Topic Segmentation

Despite recent advances in dialogue topic segmentation, existing work provides limited evidence about why individual utterances are predicted as boundaries and how local explanations depend on the selection strategy and perturbation protocol. Using fixed checkpoints of 3LHSeg, a hierarchical dialogue topic segmentation...

Fayçal Nouar, H. Belhadef · 0 citations
Jul 2026

Candidate Attended Dialogue State Tracking Using BERT

This paper presents a novel scalable framework for multi-domain dialogue state tracking that leverages the pretrained BERT model to achieve zero-shot generalization, making it easy to quickly adapt to new domains without additional training.

Junyuan Zheng, O. Salvi, John Chan · 0 citations
#natural language process... Preprint Sep 2026

Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation

In persona-based dialogue generation (PDG), LLMs often overuse persona attributes by incorporating them regardless of dialogue context, resulting in unnatural responses. Despite its practical significance, the underlying causes remain unexplored, with no method to mitigate this problem or metric to assess the appropria...

Jongkyung Shin, Inkyu Lee, Chiehyeon Lim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.