A multi-agent refinement process that iteratively improves LLM-driven counter-narratives for persuasiveness, emotional engagement, and shareability is introduced, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation.
Abstract
The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a promising way to reduce those risks, yet their effectiveness depends on rhetorical and stylistic choices that remain poorly understood. We present a multi-stage agent-based framework for generating, refining, and evaluating CNs, applied to pro-Russian hate and misinformation narratives on the war with Ukraine and adaptable to other domains. A pilot experiment with human evaluators identifies effective technique style pairings, such as repetition with emotional framing enhancing persuasiveness. Building on these insights, we introduce a multi-agent refinement process that iteratively improves CNs for persuasiveness, emotional engagement, and shareability. After human validation confirmed improvement, an automated safety analysis shows that our refined CNs match or improve on expert-written counterspeech. A simulated experiment then shows that they reduce the perceived strength of pro-Russian narratives and consistently outperform a vanilla LLM baseline, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation. Code and data accompanying this work are publicly available at https://github.com/carmelkron/inlg2026-counter-narratives.
Social media platforms have increasingly become conduits for the dissemination of hate speech targeting the LGBTQ+ community, significantly undermining their mental health and overall well-being. This phenomenon presents a critical challenge within the field of Natural Language Processing (NLP). While hate speech d...
Amritha Prabakaran, Shunmuga Priya Muthusamy Chinnan, Bharathi Raja Chakravarthi· Social Network Analysis and...· 0 citations
This study focuses on the initial stages of a broader design project aimed at developing a GenAI mediator for consensus-oriented online deliberation, and derives design knowledge, articulated through design requirements, for emotion-aware GenAI mediation.
Antoine Danthine, Anthony Simonofski· EGOV-CeDEM-ePart 2026· 0 citations
BluePRINT is introduced, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module, and Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory.
Si-Yu Chen, Hao-Ran Wang, Xiaojian Li et al.· 0 citations
Organizations have long treated clarity, timeliness and candour as the pillars of effective feedback. Yet even well-structured feedback routinely fails to produce learning. This article argues that the dominant delivery-focused framework rests on an incomplete assumption. We introduce psychological interpretability: th...
Rahul K. Shukla, Sunil Kumar Sarangi· Management and Labour Studie...· 0 citations
Large language models (LLMs) are increasingly mediating public conversations through chatbots, search assistants, and content generation tools. In the context of grey zone competition, which involves hostile activity below the threshold of open conflict, LLM systems can be repurposed to amplify cognitive warfare by inf...
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.