Skip to content
Review Open access

Benchmark Research on Safety Value Alignment Evaluation of Open-Domain Dialogue Systems based on NLP

Aug 2026 · Scientific Journal of Intelligent Systems Research · 0 citations · 21 references

TL;DR

This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.

Abstract

Large language models sometimes behave in puzzling ways. They pass various safety tests, yet in multi-turn dialogues, a few carefully crafted sentences can lead them astray into making dangerous judgments. We call this "extreme value alignment failure." A review of recent research reveals an awkward situation: attack, defense, evaluation, and theoretical studies operate in isolation, with little connection among them. This fragmentation results in repeated extreme risks that remain unresolved. This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced...

Omar Sheta, Rinku Deuja, Hadi Masoudi et al. · 0 citations
Review Open access 2026

Conversational Grounding in Large Language Models: Evaluation Methods, Challenges and Future Directions

This paper surveys how conversational grounding is evaluated in task-oriented dialogue in the current era of LLMs and focuses on how conversational grounding is modelled explicitly—using dialogue acts and by modelling the participant mental state.

Michelle Elizabeth, Gwénolé Lecorvé, L. Rojas-Barahona et al. · 0 citations
Open access 2026

Building Task-Oriented Dialogue Systems via Instruction Guidance without Annotated Data

Task-oriented dialogue (TOD) systems conventionally rely on supervised fine-tuning over large datasets, an approach that is both resource-intensive and difficult to generalize across domains. We investigate whether large language models (LLMs) can serve as effective TOD agents without any fine-tuning, relying solely on...

Henry Gao, Jinho D. Choi · 0 citations
#artificial intelligence Preprint Sep 2026

Clarification Is Not Correction: LLMs Fail to Let Go

Research for interactive LLMs from retaining more context toward preserving uncertainty is redirected from retaining more context toward preserving uncertainty, to motivate uncertainty-preserving state management.

Jian-Zhe Lin, Xiao-Lin Li, Fei Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence

Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessi...

Yu-Zhe Zhang, Wei-Jie Zhu, Hao-Lin Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.