Skip to content

Candidate Attended Dialogue State Tracking Using BERT

Jul 2026 · arXiv.org · Vol abs/2607.16021 · 0 citations · 27 references
Computer Science

TL;DR

This paper presents a novel scalable framework for multi-domain dialogue state tracking that leverages the pretrained BERT model to achieve zero-shot generalization, making it easy to quickly adapt to new domains without additional training.

Abstract

Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimates the user belief or dialogue state, which is used as input for downstream modules to predict system actions and generate responses. The increasingly popular dialogue system applications like Google Assistant, Siri and Alexa need to support a large number of services and APIs, resulting in growing attention to the scalability of such systems. Especially for some domains with little or no training data, the capability of transferring existing knowledge of other domains is highly desired. In this paper, we present a novel scalable framework for multi-domain dialogue state tracking. The proposed system leverages the pretrained BERT model to achieve zero-shot generalization, making it easy to quickly adapt to new domains without additional training. The performance of our model is evaluated on recently released schema-based dialogue (SGD) dataset, showing significant improvement compared to previous baseline.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations

This work introduces IRWOZ 2.0, which addresses limitations through large language model (LLM) enhanced generation (Mistral/Claude-3.5) and quality refinements and expands the dataset to 390 dialogues across 4 industrial domains, featuring manual corrections and automated typo removal.

Chen Li, Dimitrios Chrysostomou · 0 citations
Jul 2026

Latent-IM: Latent Interaction Management for Speech LLMs

Latent-IM is introduced, an internal dialogue-management framework that provides a general interface for choosing and deploying conversational moves under different objectives and is used to reproduce human move choices, improving average end-to-end move accuracy by 12.5 points over the unsteered backbone while perform...

Adar Avsian, Atahan Dokme, Tony Woo et al. · 0 citations
Open access Jul 2026

Evaluating reinforcement learning from human feedback for task‐oriented dialogue systems

Reinforcement learning from human feedback (RLHF) has shown strong potential for aligning language models, but its role in task‐oriented dialogue (TOD) remains unclear. In TOD, models are typically trained with local turn‐level supervision, while system behavior is evaluated through broader interaction‐level properti...

Hyeok‐Min Gwon, Yohan Lee, Jin-Xia Huang et al. · 0 citations
#natural language process... Preprint Aug 2026

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level,...

Kaiqi Fu, Rime Wen, Altman Lin et al. · 1 citation

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialo...

Neda Jamshidi, Kamyar Zeinalipour, F. Akbari et al. · 0 citations
Open access Aug 2026

Persona-centric Metamorphic Relation Guided Robustness Evaluation for Multi-turn Dialogue Modeling

This work discovers persona-centric metamorphic relations to infer test samples from annotated data, without additional annotation cost, and evaluates the robustness of personalized dialogue models regarding persona consistency, revealing that prompt learning is more robust than training from scratch and fine-tuning.

Lin Li, Xiaohua Wu, Yanbing Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.