Easy to Read, Easy to Trust: How Processing Fluency in LLM Explanations Drives Over-Reliance in Hate Speech Moderation
Abstract
This study examines the challenges human moderators face in AI-assisted detection of online hate speech, particularly when discriminatory intent is conveyed through implicit or nuanced language. Following the Judge-Advisor System (JAS) paradigm, we evaluate how different types of Large Language Model (LLM)-generated explanations support human moderators in hate speech detection. We hypothesized that natural language explanations would offer more interpretable rationales, helping moderators better assess the accuracy of AI recommendations and make informed decisions that reflect appropriate reliance on AI. Contrary to this expectation, our results show that explanation form does not affect decision accuracy but significantly reshapes reliance: fluent natural language explanations heighten moderators’ over-reliance on the AI—leading them to follow it even when it errs—whereas feature-based explanations promote scrutiny and self-reliance. We interpret this phenomenon through the lens of processing fluency within the meta-reasoning framework. Drawing on these cognitive science insights and the observed variation in how moderators process different explanation types, we provide practical design suggestions for AI-assisted moderation tools and call for greater cross-disciplinary engagement in future HCI research on human-AI decision-making.