Skip to content
Open access

Leveraging Large Language Models to Detect and Revise Unsafe Responses in Context-Sensitive Dialogues

Aug 2026 · WOCHAT2026: Workshop on Chatbots and Agentic Technologies · 0 citations · 27 references

TL;DR

This work proposes a pipeline that leverage LLMs as safety detector, editor and evaluator to mitigate undesired behaviour in human-computer dialogues and shows reduction in the unsafe dialogues after revision.

Abstract

Large Language Models (LLMs) excel at tasks like classification, summarisation, question answering among others, with performance comparable to humans. Despite these capabilities, leveraging LLMs to transform unsafe responses in context-sensitive dialogues is underexplored. In this work, we propose a pipeline that leverage LLMs as safety detector, editor and evaluator to mitigate undesired behaviour in human-computer dialogues. At the first iteration, our experimental results on two evaluation datasets show reduction in the unsafe dialogues from 47% to 13% and 48% to 2% respectively, with 82% and 92% agreement between the safety detector and evaluator after revision. Human evaluation of randomly sampled dialogues demonstrates reduction in unsafe responses after revision. Additionally, the revision LLM (editor) exhibits a higher proportion of refusals without compromising fluency and coherence of the revised dialogues.

Read PDF

Similar papers

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

Experiments with six LLAMA and MISTRAL models show that LoRA substantially improves dialogue generation while requiring only 25-50% of the training data to recover over 90% of the final performance gains, and zero-shot evaluation with frontier LLMs shows that TalkFa remains a challenging benchmark.

Neda Jamshidi, Kamyar Zeinalipour, F. Akbari et al. · 0 citations
#natural language process... Preprint Sep 2026

Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation

In persona-based dialogue generation (PDG), LLMs often overuse persona attributes by incorporating them regardless of dialogue context, resulting in unnatural responses. Despite its practical significance, the underlying causes remain unexplored, with no method to mitigate this problem or metric to assess the appropria...

Jongkyung Shin, Inkyu Lee, Chiehyeon Lim · 0 citations
Open access Aug 2026

Context-Aware Large Language Model for Customer Support Chatbots

The results show that the RAG architecture provides a scalable alternative for creating precise, contextually grounded conversational agents, thereby mitigating some of the main drawbacks of LLMs.

Rabia Shabbir, K. Talpur, Shakeel Ahmad · 0 citations
Open access Aug 2026

Bridging context gaps in low-resource language chatbots through multilevel attention and hybrid embedding approaches

A multilevel-attention and hybrid-embedding framework that integrates FastText subword representations with multilingual BERT with potential applicability to other African languages is proposed to improve semantic understanding and context retention in conversational agents for Igbo.

G. C. Uzoaru, I. Ayogu, J. N. Odii et al. · 0 citations
Open access Jul 2026

Xbot: a GPT-based chatbot with transparent and empathetic behaviour

Experimental comparisons with GPT-4o vanilla across three roles, evaluated through an ablation study and a multi-evaluator panel combining LLM-based and human judges, consistently rank XBot as the best performing system, demonstrating superior empathy, role stability and conversational depth, while GPT-4o vanilla exhib...

Luciano Caroprese, Ester Zumpano, M. Aracne et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.