Skip to content
Preprint

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

Aug 2026 · 0 citations · 53 references
Computer Science

Abstract

Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.

View source

Similar papers

Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses ac...

Katherine Van Koevering, A. Field · 0 citations
Book Open access Oct 2026

Judging Turn-Taking Through a Monocultural Lens: A Parallel-Corpus Probe of Cultural Bias in LLMs

Turn-taking and backchanneling are governed by culture-specific display rules: the same listener behavior that signals attentive engagement in one language community can read as interruption or disengagement in another. As large language models (LLMs), often accessed as multimodal systems, are increasingly deployed as...

Nur Keleşoğlu · 1 citation
#artificial intelligence Preprint Aug 2026

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

A multilingual German-English benchmark dataset that combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer is introduced, showing that language models reproduce anti-queer stereotypes, with variation across identities and models.

M. Morch, Daniel Braun · 0 citations
#artificial intelligence Preprint Aug 2026

How Identity and Opinion Shape Political Sycophancy in LLMs

A framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels) is introduced, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.

Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen et al. · 0 citations
#natural language process... Preprint Sep 2026

Framing the Narrative: Ideological Mimicry in Large Language Models

Large language models (LLMs) are increasingly used to answer questions about politically contentious issues, yet evaluations typically treat a model's stance as a relatively stable property. Real users, however, communicate political signals through their terminology, assumptions, and personal context. We investigate w...

Olivia Macmillan-Scott, Michael Jacobs, Nils W. Metternich et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.