The results of this research demonstrate that multimodal fusion and cross-language alignment mechanisms offer effective approaches to improve the accuracy and reliability of multimodal translation systems.
Abstract
Translation across multiple forms and different languages shows limitations from problems with combining features and problems with moving meaning between languages, and these problems affect how the approach performs in settings that involve various conditions. This study aims to design a unified multimodal Transformer architecture that strengthens cross-linguistic semantic alignment and improves translation stability under heterogeneous input conditions. This study conducted initial processing that combined image and text data, allowing separate encoding of features at higher levels. It then used attention across different forms to combine meaning from multiple sources, producing a representation that showed unity. The approach also introduced space for meaning that multiple languages share at a level beyond surface forms, thereby improving the stability of the mapping between languages in the results. Training without modality and applying noise-enhancement techniques together improved this model’s robustness to incomplete and degraded input conditions that approximate practical deployment scenarios.The model was trained and evaluated on a structured English-German bilingual vision-language dataset composed of aligned image-text pairs collected under controlled experimental conditions.Experimental evaluation in fusion mode yielded BLEU scores of 39.2, METEOR scores of 31.7, and CIDEr scores of 113.8, indicating stronger semantic consistency and improved bilingual generation quality. The results of this research demonstrate that multimodal fusion and cross-language alignment mechanisms offer effective approaches to improve the accuracy and reliability of multimodal translation systems. The findings remain constrained by dataset scale and language coverage, and future research will extend validation across broader linguistic domains and more diverse multimodal scenarios.
Machine translation, as a core approach in natural language processing, plays a crucial role in promoting cross-lingual communication. This study proposes an English-to-target language translation model that integrates enhanced back-translation with translation memory, aiming to address the issues of low training effic...
Yang Du· HighTech and Innovation Jour...· 0 citations
This paper introduces a multimodal collaborative representation learning algorithm (MCRA-Net) to tackle challenges in English translation, including cultural differences, context dependency, terminological accuracy, and polysemy resolution. The algorithm integrates three types of heterogeneous data-text, image and voic...
Xia Wu· International Conference on...· 0 citations
For a long time, there has been a lack of effective cultural semantic conversion and automatic detection methods for term inconsistency in the English Chinese translation of cultural heritage terminology. This paper constructs an automatic error detection model for cultural heritage terminology translation based on Bid...
A Contrastive Learning-based Chinese-English Scientific Translation Quality Evaluation model (C-TQE), which provides an effective solution for large-scale scientific translation quality assessment and facilitates the accurate international communication of multidisciplinary engineering research, including electromagnet...
It is shown that translation is even more modular than previously assumed and that the output language production in translation processes is actually further separable into a syntax and a surface language process.
M. Sonkin, Tanja Baeumel, Daniil Gurgurov et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.