Skip to content
Open access

Large Language Models for Japanese–Croatian Translation: Human Evaluation and Macroeconomic Implications

Jul 2026 · Acta Linguistica Asiatica · Vol 16, pp. 237-261 · 0 citations

Abstract

Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.

Read PDF

Similar papers

Open access Jul 2026

Croatian Language in the Transition from Neural Machine Translation to Large Language Models

Machine translation (MT) technologies are currently undergoing a paradigm shift, transitioning from specialized Neural Machine Translation (NMT) frameworks to the broader capabilities of Large Language Models (LLMs). This paper examines the current standing of the Croatian language within this technological evolution. While bilingual NMT models often exhibit high precision, multilingual NMT leverage transfer learning to enhance performance for low–resource language pairs, but with lower performance for high–resource ones. Conversely, LLMs—whether general–purpose or fine–tuned for translation— offer superior multilingual proficiency and context awareness. Unlike NMT, LLMs can process extended discourse, such as full paragraphs or documents, leading to significant improvements in coreference resolution and gender agreement. Despite the substantial computational requirements of LLMs, recent optimization techniques allow for smaller, more efficient versions that maintain high output quality. This study evaluates the performance of various NMT and LLM architectures specifically for Croatian from/to English and Spanish using several automatic quality evaluation metrics. The findings demonstrate that open–source models can achieve, and occasionally surpass, the quality of Google Translate, a widely used commercial NMT system. Furthermore, while our evaluation focuses on this specific language triad, the multilingual nature of the analysed systems suggests that open–source models provide high–quality translation capabilities for Croatian across dozens, if not hundreds, of language pairs.

A. Oliver, Sergi Álvarez–Vidal · 1 citation
Preprint Aug 2026

Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs

J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.

Florian Braun · 1 citation
Preprint Aug 2026

Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese

Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers'language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.

Naho Orita, Hayato Ogawa, Daisuke Kawahara · 0 citations
Open access Jul 2026

Evaluating the limits of machine translation for poetry: a multidimensional framework

The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.

Beatriz Ribeiro Borges, P. H. R. Gabriel, E. Faria · 0 citations
Aug 2026

Taming the stochastic parrot

This article proposes a Translation Studies–oriented approach to understanding and managing variability in machine translation outputs generated by large language models (LLMs). Drawing on the metaphor of the “stochastic parrot,” the study introduces the concept of temperature as a means for controlling stochasticity in LLM-based translation. Through a practical and replicable experiment conducted in Google Colab, technical texts are translated from English to Spanish under varying temperature conditions. While the dataset is intentionally limited, the study’s primary contribution lies in establishing a replicable methodological pathway rather than in producing generalizable quantitative results. By combining computational experimentation with reflection on concepts relevant to translation theory, the study can inform both future research and practical approaches.

Celia Rico Pérez · 0 citations
Review Open access Aug 2026

Evaluating Terminological Consistency in AI-Generated English–Arabic Political Translation

Machine translation has become more fluent and contextually accurate with recent advances in artificial intelligence. However, terminological consistency has been underexplored, particularly in political and electoral discourse where lexical repetition and conceptual precision are critical for cohesion and clarity. This study investigates terminological consistency in AI-generated English–Arabic political translations produced by ChatGPT and Google Gemini Advanced. The translations were generated and analyzed between January and June 2026 using the systems’ default settings to ensure comparability and avoid potential variations resulting from user-configured parameters. The study employs a mixed-methods corpus-based approach. The study analyzes 30 political and electoral texts with 60 recurring key terms. Quantitative analysis measures the degree of consistency in the form of stability percentages, and qualitative analysis studies lexical variation and its effect on discourse cohesion and clarity. The adequacy and consistency of the translation were checked against a reference translation based on the United Nations Development Programme (UNDP) Arabic Lexicon of Electoral Terminology. The results indicate that ChatGPT achieved higher terminological consistency than Google Gemini. ChatGPT’s lexical equivalents for repeated political and electoral terms were more stable than its Gemini counterpart, which showed more lexical variation, especially in context-sensitive terms such as campaign, electoral law, and judicial review. The study concludes that terminological consistency should be considered as a separate dimension of translation quality and emphasizes the importance of terminology control and human post-editing in AI-assisted political translation

Osama Bala · 0 citations