Skip to content
Preprint

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

Aug 2026 · 0 citations · 71 references
Computer Science

TL;DR

A per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models are released.

Abstract

Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models.

View source

Similar papers

Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses ac...

Katherine Van Koevering, A. Field · 0 citations
Preprint Aug 2026

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

The findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone, and that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge.

Mena Attia, Mona T. Diab, Thamar Solorio · 0 citations
#natural language process... Preprint Oct 2026

Language Unalignability: Why Some Concepts Resist Cross-Cultural Benchmark Evaluation

Current evaluation of multilingual Large Language Models (LLMs) rests on an implicit Translation-Isomorphism Assumption (TIA): that semantic structures across languages are congruent and mutually mappable without loss of information. We argue that this assumption is not merely violated in practice, but ill-posed in pri...

Shu-Kai Hsieh, Da-Chen Lian · 0 citations
Preprint Aug 2026

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obs...

Yuanjun Feng, Tanzhou Liu, S. Feuerriegel et al. · 0 citations

Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model

Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, w...

Fathin Difa Robbani · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.