Aug 2026· Scientific Journal of Intelligent Systems Research· 0 citations· 21 references
TL;DR
This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.
Abstract
Large language models sometimes behave in puzzling ways. They pass various safety tests, yet in multi-turn dialogues, a few carefully crafted sentences can lead them astray into making dangerous judgments. We call this "extreme value alignment failure." A review of recent research reveals an awkward situation: attack, defense, evaluation, and theoretical studies operate in isolation, with little connection among them. This fragmentation results in repeated extreme risks that remain unresolved. This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced...
Omar Sheta, Rinku Deuja, Hadi Masoudi et al.· 0 citations
This paper surveys how conversational grounding is evaluated in task-oriented dialogue in the current era of LLMs and focuses on how conversational grounding is modelled explicitly—using dialogue acts and by modelling the participant mental state.
Michelle Elizabeth, Gwénolé Lecorvé, L. Rojas-Barahona et al.· SIGDIAL Conferences· 0 citations
Task-oriented dialogue (TOD) systems conventionally rely on supervised fine-tuning over large datasets, an approach that is both resource-intensive and difficult to generalize across domains. We investigate whether large language models (LLMs) can serve as effective TOD agents without any fine-tuning, relying solely on...
Henry Gao, Jinho D. Choi· SIGDIAL Conferences· 0 citations
Research for interactive LLMs from retaining more context toward preserving uncertainty is redirected from retaining more context toward preserving uncertainty, to motivate uncertainty-preserving state management.
Jian-Zhe Lin, Xiao-Lin Li, Fei Wang et al.· 0 citations
Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessi...
Yu-Zhe Zhang, Wei-Jie Zhu, Hao-Lin Yang et al.· 0 citations
The results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.
Nan Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.