Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.
Machine translation (MT) technologies are currently undergoing a paradigm shift, transitioning
from specialized Neural Machine Translation (NMT) frameworks to the broader capabilities of
Large Language Models (LLMs). This paper examines the current standing of the Croatian language
within this technological evolution.
While bilingual NMT models often exhibit high precision, multilingual NMT leverage transfer
learning to enhance performance for low–resource language pairs, but with lower performance for
high–resource ones. Conversely, LLMs—whether general–purpose or fine–tuned for translation—
offer superior multilingual proficiency and context awareness. Unlike NMT, LLMs can process extended discourse, such as full paragraphs or documents, leading to significant improvements in
coreference resolution and gender agreement. Despite the substantial computational requirements
of LLMs, recent optimization techniques allow for smaller, more efficient versions that maintain
high output quality.
This study evaluates the performance of various NMT and LLM architectures specifically for
Croatian from/to English and Spanish using several automatic quality evaluation metrics. The findings demonstrate that open–source models can achieve, and occasionally surpass, the quality of
Google Translate, a widely used commercial NMT system. Furthermore, while our evaluation focuses on this specific language triad, the multilingual nature of the analysed systems suggests that
open–source models provide high–quality translation capabilities for Croatian across dozens, if not
hundreds, of language pairs.
A. Oliver, Sergi Álvarez–Vidal· Suvremena Lingvistika· 1 citation
J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.
Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers'language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.
The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.
Beatriz Ribeiro Borges, P. H. R. Gabriel, E. Faria· International Journal of Dat...· 0 citations
This article proposes a Translation Studies–oriented approach to understanding and managing variability in machine
translation outputs generated by large language models (LLMs). Drawing on the metaphor of the “stochastic parrot,” the study
introduces the concept of
temperature
as a means for controlling stochasticity in LLM-based translation. Through
a practical and replicable experiment conducted in Google Colab, technical texts are translated from English to Spanish under
varying temperature conditions. While the dataset is intentionally limited, the study’s primary contribution lies in establishing
a replicable methodological pathway rather than in producing generalizable quantitative results. By combining computational
experimentation with reflection on concepts relevant to translation theory, the study can inform both future research and
practical approaches.
Celia Rico Pérez· Digital Translation· 0 citations
Machine translation has become more fluent and contextually accurate with recent advances in artificial intelligence. However, terminological consistency has been underexplored, particularly in political and electoral discourse where lexical repetition and conceptual precision are critical for cohesion and clarity. This study investigates terminological consistency in AI-generated English–Arabic political translations produced by ChatGPT and Google Gemini Advanced. The translations were generated and analyzed between January and June 2026 using the systems’ default settings to ensure comparability and avoid potential variations resulting from user-configured parameters. The study employs a mixed-methods corpus-based approach. The study analyzes 30 political and electoral texts with 60 recurring key terms. Quantitative analysis measures the degree of consistency in the form of stability percentages, and qualitative analysis studies lexical variation and its effect on discourse cohesion and clarity. The adequacy and consistency of the translation were checked against a reference translation based on the United Nations Development Programme (UNDP) Arabic Lexicon of Electoral Terminology. The results indicate that ChatGPT achieved higher terminological consistency than Google Gemini. ChatGPT’s lexical equivalents for repeated political and electoral terms were more stable than its Gemini counterpart, which showed more lexical variation, especially in context-sensitive terms such as campaign, electoral law, and judicial review. The study concludes that terminological consistency should be considered as a separate dimension of translation quality and emphasizes the importance of terminology control and human post-editing in AI-assisted political translation
Osama Bala· (Faculty of Arts Journal) مج...· 0 citations