Judging Turn-Taking Through a Monocultural Lens: A Parallel-Corpus Probe of Cultural Bias in LLMs
Abstract
Turn-taking and backchanneling are governed by culture-specific display rules: the same listener behavior that signals attentive engagement in one language community can read as interruption or disengagement in another. As large language models (LLMs), often accessed as multimodal systems, are increasingly deployed as evaluators of conversational behavior, it is unclear whether their judgments encode such culture-specific norms or default to a single, implicitly English-language, interactional standard. We introduce a lightweight probing protocol that isolates cultural bias using parallel dialogue: identical conversational content aligned across four languages (English, German, Italian, Chinese) from the XDailyDialog corpus. Holding content fixed, we elicit LLM judgments of turn-appropriateness, response naturalness, and speaker engagement, and measure how these judgments shift with the language and stated cultural framing of the interlocutors. On a text-only open-weight model, we find that identical dialogues are rated significantly more appropriate and more engaged in German and Italian than in English (drifts up to + 0.49 on a 5-point scale), and that merely relabeling a dialogue’s culture, without changing its words, can lower its rating by more than a point, isolating the effect of the cultural label from that of the content. The drift is stable across repeated sampling but differs in direction across models, indicating that current LLMs assess turn-taking through a monocultural, model-specific lens.