Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.
SocraticTrap-CS is introduced, a publicly available benchmark that probes the capacity of open-weight LLMs to generate strategic misconceptions on demand and reframes the evaluation of educational LLMs around pedagogical trustworthiness rather than factual correctness alone.
Marijela Miličević, Mia Rovis, Ratomir Karlović et al.· Information· 0 citations