Skip to content
Preprint

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

The findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone, and that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge.

Abstract

Figurative language is deeply culturally embedded; fluent use requires not just linguistic competence but cultural immersion. We ask whether LLMs can learn this link: does fine-tuning on cultural data improve figurative language understanding, and vice versa? We conduct a systematic study across four models (ALLaM-7B, Fanar-1-9B, Qwen3-8B, Llama-3.1-8B) and six Arabic datasets spanning cultural commonsense, proverbs, and poetry across diverse dialects and regions. Fine-tuning on poetry improves idiom comprehension (+2.33%, p<0.05), a gain our ArabicMMLU control does not reproduce, indicating that it stems from figurative content rather than Arabic language adaptation and pointing to a sensitivity to non-literal meaning that transfers across figurative types. Cultural fine-tuning, by contrast, lowers proverb-interpretation accuracy in both Arabic-centric models. Transfer between the two domains is otherwise indistinguishable from noise, with Arabic models frequently regressing after fine-tuning, suggesting prior saturation of relevant knowledge, while multilingual models show greater adaptation headroom. Error analysis further reveals that fine-tuning reinforces experiential cultural knowledge while destabilizing historically grounded factual knowledge. Our findings suggest that the relationship between culture and figurative language, though conceptually natural, is not straightforwardly captured through fine-tuning alone.

View source

Similar papers

Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
Open access Aug 2026

Cross-Cultural Scenario Benchmark: Evaluating LLMs’ Cross-Cultural Understanding

Cross-cultural reasoning and alignment have been identified as key weaknesses of large language models (LLMs), but the architectural or cognitive features underlying these failures have not been adequately examined. In addition, previous studies rely almost exclusively on datasets and benchmarks constructed under the W...

Meng-Xi Guo, Lei-Ming Gao, W. Zeng et al. · 0 citations
#natural language process... Preprint Aug 2026

Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs

This work introduces multidimensional annotation framework for proverbs that characterizes proverbs through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture-Specific, and shows that combining diverse figurative forms yields the strongest overall performance.

Rama Alomair, Remas Alsubaie, Walaa Saifalislam et al. · 0 citations
Preprint Aug 2026

Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs

A per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models are released.

Iaroslav Chelombitko, Ekaterina Chelombitko, Mika K. Hämäläinen · 0 citations
Open access Jul 2026

Reframing collocational errors as cultural collocational transfer: Insights from a Thai EFL learner corpus

This study investigates how culturally grounded conceptualizations shape English collocational usage in learner writing. Drawing on a 4.6-million-word corpus of Thai English as a foreign language (EFL) academic texts, collocations were extracted using corpus-driven association measures, including Mutual Information (MI...

Songtham Vongvirulh, A. Khamkhien · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.