Skip to content
Open access

Quantification of CITTA Chinese Translation Interpretation Differentiation based on BEiT-RoBERTa Model

Aug 2026 · WSEAS Transactions on Computer Research · pp. 479 · 0 citations · 10 references

TL;DR

Research has confirmed that the BEiT-RoBERTa model quantifies CITTA translation variations but poorly captures philosophical depth and cultural subtleties.

Abstract

This study analyzes varied Chinese renderings of the Buddhist core concept CITTA via quantitative modeling to expose cross-cultural translation discrepancies. It constructs a BEiT-RoBERTa fusion model combining BEiT visual feature extraction and RoBERTa linguistic encoding. Based on a self-built CITTA Chinese translation text dataset containing ancient and modern translations, experiments were conducted from four dimensions: vocabulary recognition, model comparison, semantic coherence, and multi-context adaptation. Performance tests were conducted using indicators such as character error rate (CER), syllable recognition accuracy (SRA), line recognition accuracy (LRA), and manual evaluation. Results show the syllable-based model performs best: CER=0.167, SRA=0.850, LRA=0.791, surpassing CRNN and SRN. The overall accuracy of sentence semantic coherence judgment is 92.68%, and the translation effect in different contexts is between professional translators and non professional enthusiasts, with stable accuracy and fluency. Research has confirmed that the BEiT-RoBERTa model quantifies CITTA translation variations but poorly captures philosophical depth and cultural subtleties.

Read PDF

Similar papers

Open access Jul 2026

Evaluating the limits of machine translation for poetry: a multidimensional framework

The results show that LLMs and Google Translate consistently outperform specialized MT systems in terms of fluency, meaning preservation, and lexical-thematic alignment.

Beatriz Ribeiro Borges, P. H. R. Gabriel, E. Faria · 0 citations
Open access Jul 2026

Cognitive Misalignment and Categorical Reconstruction of Qi in English Translation

The translation of culture-specific items is difficult not merely because of lexical gaps, but because languages often organize experience through different underlying cognitive categories. This study examines the Chinese core concept qi (Gas), a radial category rooted in embodied experience and centered on the prototype of vital energy, and explains why it is frequently fragmented in English translation. A bilingual database of 49 high-frequency qi-compounds was built from the Chinese Proficiency Grading Standards for International Chinese Language Education and checked against the Modern Chinese Dictionary and Oxford Learner's Dictionaries. The static morphological comparison was then supplemented by qualitative analysis of authentic translations from medical classics, philosophy, literary theory, ethical discourse, and fiction. The results show a systematic mismatch: Chinese organizes the qi-word family through a shared morpheme and modifier-head construction, with 34 of 49 items placing qi in final position and 37 of 49 exhibiting modifier-head structure. By contrast, English renders these items through morphologically unrelated simplexes, derivatives, phrases, and one compound, thereby distributing a continuous Chinese category across discrete lexical domains such as air, strength, anger, courage, integrity, and atmosphere. Translation examples further reveal metaphorical rupture when the force, container, and flow schemas of qi are replaced by static English entities. To address this problem, the study proposes a cognitive compensation framework consisting of category transplantation, semantic compensation, category re-implantation, and context-sensitive modulation. The framework contributes to cognitive translation studies and offers practical implications for translating and teaching culturally dense Chinese key concepts in applied language-learning contexts.

Ting Zhang, Zihan Yu · 0 citations
Review Open access Aug 2026

Evaluating Terminological Consistency in AI-Generated English–Arabic Political Translation

Machine translation has become more fluent and contextually accurate with recent advances in artificial intelligence. However, terminological consistency has been underexplored, particularly in political and electoral discourse where lexical repetition and conceptual precision are critical for cohesion and clarity. This study investigates terminological consistency in AI-generated English–Arabic political translations produced by ChatGPT and Google Gemini Advanced. The translations were generated and analyzed between January and June 2026 using the systems’ default settings to ensure comparability and avoid potential variations resulting from user-configured parameters. The study employs a mixed-methods corpus-based approach. The study analyzes 30 political and electoral texts with 60 recurring key terms. Quantitative analysis measures the degree of consistency in the form of stability percentages, and qualitative analysis studies lexical variation and its effect on discourse cohesion and clarity. The adequacy and consistency of the translation were checked against a reference translation based on the United Nations Development Programme (UNDP) Arabic Lexicon of Electoral Terminology. The results indicate that ChatGPT achieved higher terminological consistency than Google Gemini. ChatGPT’s lexical equivalents for repeated political and electoral terms were more stable than its Gemini counterpart, which showed more lexical variation, especially in context-sensitive terms such as campaign, electoral law, and judicial review. The study concludes that terminological consistency should be considered as a separate dimension of translation quality and emphasizes the importance of terminology control and human post-editing in AI-assisted political translation

Osama Bala · 0 citations
Open access Jul 2026

Large Language Models for Japanese–Croatian Translation: Human Evaluation and Macroeconomic Implications

Large Language Models (LLMs) are increasingly used for translation, yet their value depends on preserving meaning rather than producing fluent output. This study evaluates seven LLMs on Japanese–Croatian translation, a low-resource, typologically distant language pair. Using rubric-based human evaluation of adequacy, fluency, terminology, and register, we compare model performance. Results show a stable ranking: qwen3 performs best, followed by phi4 and gemma3, while qwen2 performs worst. Performance differences reflect structural reconstruction, particularly argument recovery, aspectual mapping, lexical precision, and register. Qualitative analysis also reveals limited differentiation within the South Slavic continuum and pragmatic inconsistencies. Although productivity effects were not measured, improved translation adequacy may reduce post-editing and verification effort.

Ratomir Karlović, Mieta Bobanović Dasko, Irena Srdanović · 0 citations
Open access Aug 2026

Zoonym-Based Phraseology in English–Azerbaijani AI-Assisted Translation: Semantic, Affective, and Pragmatic Equivalence

Artificial-intelligence-assisted translation of culturally marked phraseology may produce semantically plausible output while failing to preserve affective valence, pragmatic force, register, or cultural symbolism. This study examines English and Azerbaijani zoonym-based phraseological units in bidirectional AI-assisted translation. A qualitative corpus-based comparative design was applied to 100 units (50 English and 50 Azerbaijani) selected from phraseological dictionaries and literary and digital sources. Translations were generated with ChatGPT (OpenAI GPT-5.5) through the official web interface between 15 and 20 July 2026. Each source item was tested three times in newly initiated sessions using a standardized prompt. Outputs were compared with reference equivalents identified in the cited lexicographic sources and verified by the author across six dimensions: semantic adequacy, idiomatic naturalness, affective equivalence, pragmatic function, cultural appropriateness, and register preservation. The qualitative case analyses illustrate successful preservation of conventional target-language equivalents in some items and literal rendering, metaphorical-image mismatch, reduced emotional expressiveness, pragmatic weakening, or loss of cultural symbolism in others. Because corpus-level frequencies and the complete item-level record are not presented in this version, the findings should be interpreted as qualitative patterns rather than statistical estimates. The study demonstrates the value of context-sensitive, linguoculturally informed evaluation of AI-assisted phraseological translation.

Aysel V. Safarova · 0 citations
Open access Aug 2026

On the Chinese Translation Strategies under Eco-translatology

Building on the theoretical approach of Eco-translatology, the paper analyses the translation strategies used by Huang Yuanshen in his Chinese translation of 《归宿》 (To the Islands) through linguistic, cultural, and communicative dimensions. It can be found that, linguistically, Huang uses a multi-dimensional strategy that involves syntactic domestication together with imagistic foreignization via Chinese paratactic constructions for the poetic features of the source text. Culturally, Huang applies a combination strategy of “domestication in the title and foreignization in the main text” in dealing with the Indigenous Australian cultural references and Christian discourses, while Taoism views about nature are indirectly incorporated into the Chinese translation. Communicatively, ritual-like sentences with repetitive rhythms are utilized for bridging the writer’s intention with the audience. The study demonstrates that Huang achieves dynamic adaptation between the source-language and target-language ecologies, enabling the original work to successfully “survive” and “grow” within the Chinese literary ecosystem.

Yu-Hang Gu, Zhu-Lin Han · 0 citations