Skip to content
Open access

Evaluating the cultural alignment of multilingual LLMs in typical Japanese workplace scenarios

Jul 2026 · PLoS ONE · Vol 21, pp. e0338524 - e0338524 · 0 citations · 30 references
Medicine

Abstract

While current evaluations of LLM cultural alignment predominantly rely on static benchmarks in Western contexts, their ability to navigate generative, high-context socio-pragmatic demands in non-Western environments remains critically underexplored. This study investigates how multilingual LLMs adapt to the Japanese workplace—a stringent stress-test environment characterized by strong high-context communication norms and rigid honorific conventions—using Hofstede’s six cultural dimensions as a heuristic framework. We evaluated five state-of-the-art LLMs (LLM-jp, Phi, Llama, Qwen, and GLM) through a large-scale crowdsourced human evaluation. Based on 1,718 valid evaluator sessions, native Japanese raters assessed model outputs to generate a holistic Japanese Workplace Cultural Alignment Score (JWCAS). To dissect the underlying communicative strategies, we paired this with a three-layer diagnostic sub-score analysis (Linguistic Form, Socio-Cultural Values, and Social Action). Our results reveal that leading multilingual models (Phi and GLM) achieved overall JWCAS scores comparable to, or significantly higher than, the native Japanese model (LLM-jp). Crucially, our sub-score analysis demonstrates that holistic evaluation metrics can obscure deep pragmatic deficits: while LLM-jp overfits to surface-level linguistic politeness (Layer 1), it shows critical weaknesses in socio-cultural values (Layer 2) and context-aware social strategies (Layer 3). In contrast, leading multilingual models demonstrate balanced competence across all layers. These findings suggest that true cultural competence requires moving beyond native linguistic mastery, highlighting the necessity of multi-dimensional diagnostic frameworks for cross-cultural AI alignment.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values

As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and comprising $12$ dimensions across national, societal, and personal levels. We construct C-Voices, the first comprehensive multilingual contrastive probe dataset for CSV, with 86,400 dilemma-based instances in six languages, each pairing a CSV-aligned action with a value-conflicting alternative. Building on the contrastive probes of C-Voices, we then propose a fine-tuning-free value vector steering method that derives value directions from hidden-state discrepancies and selectively intervenes on value-sensitive layers during inference. Experiments on six languages show that CSV-oriented preferences are model-dependent and language-sensitive, with the same dilemma eliciting divergent responses across languages. Our method achieves effective CSV steering, supports cross-lingual transfer of value vectors, and generalizes to existing FLAMES and ValuePrism.

Yue-Mei Xu, Ke-Xin Xu, Jian Zhou et al. · 0 citations
Review Open access Aug 2026

Artificial Minds, Cultural Shadows: Cultural Alignment, Identity, and Voice Across Multiple Large Language Models

Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.

Ashkan Goudarzi, Aylar Naderi Zonouz · 0 citations
Open access Aug 2026

Cross-Cultural Scenario Benchmark: Evaluating LLMs’ Cross-Cultural Understanding

Cross-cultural reasoning and alignment have been identified as key weaknesses of large language models (LLMs), but the architectural or cognitive features underlying these failures have not been adequately examined. In addition, previous studies rely almost exclusively on datasets and benchmarks constructed under the WEIRD (Western, Educated, Industrialized, Rich, and Democratic) bias. To address this data bias issue, we prepare a dataset with substantial coverage of non-WEIRD cultures and five-dimensional (W, E, I, R, and D) annotations. This dataset supports an interpretable approach to examining weaknesses in LLMs’ cross-cultural alignment. We adopt the Chinese–Foreign Cultural Differences Case Repository at Xiamen University, which contains 9342 cases across 151 countries, 6 continents, and 10 cultural domains. These cases are processed and transformed into benchmark-ready structured data through topic normalization, structured metadata cleaning, continent correction, and country-level WEIRD annotation along five dimensions. Each case is converted into a six-option cultural attribution question with five cognitive-trap distractors grounded in cognitive reasoning and pragmatic interpretation. Evaluation of 6 mainstream large language models shows that their dominant failures do not involve explicit stereotypes. Instead, 61% of all errors arise from oversimplifying complex cultural phenomena or applying familiar cultural frames. The proportion of errors that explain specific cultural conflicts through seemingly universal value frames increases from 11% at the low-WEIRD end to 20% at the high-WEIRD end of the dataset. These results suggest that WEIRD data bias reflects both the underrepresentation of low-WEIRD cultures and the overactivation of dominant value frames in high-WEIRD contexts.

Meng-Xi Guo, Lei-Ming Gao, W. Zeng et al. · 0 citations
#natural language process... Preprint Aug 2026

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

CultureConverse is introduced, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains and performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks.

Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan et al. · 0 citations
Open access Jul 2026

Prompting for Pragmatics: Improving the Cultural Sensitivity of LLM Translations for Business Emails

Large language models (LLMs) are rapidly diffusing into multilingual workplaces, where managers and employees rely on them to translate everyday communication. While LLMs can now generate translations with high lexical accuracy, it remains unclear whether they can produce translations that are culturally appropriate for the target audience—an issue central to collaboration in multinational corporations. We analyze the cultural sensitivity of English-to-Japanese LLM translations of workplace emails across three prompting strategies: (1) naïve “just translate” prompts, (2) audience-targeted prompts specifying the recipient’s cultural background and workplace role, and (3) instructional prompts providing explicit guidance on Japanese communication norms. Using a convergent mixed-methods design, Study 1 applies a cross-cultural pragmatics scheme to compare culture-specific language patterns in source texts and translations; Study 2 asks Japanese native speakers to evaluate each translation for perceived appropriateness and intention to comply. We find that naïve prompts yield only limited adaptation toward Japanese language patterns. Audience-targeted prompts produce meaningful improvements, and instructional prompts generate the strongest textual adaptation. However, recipient evaluations reveal a threshold effect: although both culturally informed prompts significantly outperform the naïve baseline, instructional prompts do not produce significantly higher appropriateness or compliance ratings than audience-targeted prompts. Theoretically, we advance language-sensitive international management research by applying cross-cultural pragmatics to LLM-mediated workplace translation and by providing a replicable way to assess the cultural sensitivity of translated workplace communication through both pragmatic adaptation of language patterns and recipient evaluations. We also theorize LLMs as new translating agents in multilingual MNC communication. Practically, our findings show that lightweight audience-targeted prompting can already yield meaningful improvements in the cultural sensitivity of LLM translations.

H. Tenzer, O. Abidi, S. Feuerriegel · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.