SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue
SDARE-Bench is introduced, the first scenario-based benchmark evaluating both stigma detection and open-ended response generation in LLMs, comprising 1,138 dyadic queries and 1,388 group dialogue and identifies stigma response as a recurring LLM safety vulnerability, especially in socially complex conversational contexts.