Skip to content
Open access

Large-Language-Model-Enhanced Adaptive Optimization for Data-Limited Multi-Domain Machine Translation

2026 · IEEE Access · Vol 14, pp. 120331-120348 · 0 citations · 43 references
Computer Science

TL;DR

DKA-LLM-MT is proposed, a large-language-model-enhanced adaptive optimization framework for data-limited multi-domain machine translation that provides an effective and reliable solution for domain-sensitive machine translation under limited bilingual supervision and offers practical support for bilingual reading, specialized translation assistance, and domain-oriented language learning.

Abstract

Data-limited multi-domain machine translation remains challenging because parallel corpora are scarce in specialized domains, domain terminology is highly constrained, and large language models may generate fluent but unfaithful translations. Direct prompting or ordinary fine-tuning is therefore insufficient for domain-sensitive translation scenarios. To address these issues, this paper proposes DKA-LLM-MT, a large-language-model-enhanced adaptive optimization framework for data-limited multi-domain machine translation. The framework follows a data–model–reliability design. First, a domain-knowledge-constrained data augmentation strategy generates pseudo-parallel corpora under terminology, semantic consistency, and domain-style constraints. Second, a retrieval-augmented parameter-efficient adaptation mechanism integrates domain memory retrieval, lightweight LoRA adapters, and dynamic domain routing. Third, a reliability-aware optimization mechanism incorporates semantic fidelity, terminology consistency, and hallucination risk into both training-time data selection and inference-time candidate reranking. Experiments are conducted on five public data-limited domain translation benchmarks covering medical, legal, technical, news, and spoken-style texts. The proposed method achieves an average BLEU of 36.18, chrF of 62.14, COMET of 0.816, and TER of 40.62, consistently outperforming strong neural, multilingual, and LLM-based baselines. Additional matched-backbone and same-language-pair analyses are included to separate the effect of domain adaptation from language-pair variation. Reliability evaluation further shows that DKA-LLM-MT improves terminology accuracy to 89.6% and reduces hallucination rate to 3.2%. The proposed framework provides an effective and reliable solution for domain-sensitive machine translation under limited bilingual supervision and offers practical support for bilingual reading, specialized translation assistance, and domain-oriented language learning.

Read PDF

Similar papers

Preprint Aug 2026

PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation

PAMT is proposed, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning that improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.

Yongshi Ye, Biao Fu, Chongxuan Huang et al. · 0 citations
Open access Jul 2026

Semi-Supervised Marginal Likelihood Training with Curriculum-Guided Rewriting for Low-Resource Machine Translation

Large language models continue to face challenges in translating low-resource languages with scarce parallel data. This study investigates how to fine-tune them effectively using target-side monolingual data. Existing approaches—dominated by back-translation and recent LLM-based rewriting—remain limited by noisy synthe...

Wenjie Yu, Zhiqiang Yu, Zuo Jiang et al. · 0 citations
Jul 2026

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

Mixture-of-Translators (MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM, is proposed, demonstrating scalable KV cache reuse across heterogeneous LLMs.

Jin-woo Lee, Minkyung Song, Junghyun Oh et al. · 4 citations
Open access Aug 2026

A Data-Efficient Multilingual Neural Machine Translation Model for Low-Resource Indic Languages

The effectiveness of multilingual transfer learning in low-resource settings is demonstrated by the fine-tuned Multilingual Bidirectional and Auto-Regressive Transformer-50 model, significantly outperforming the pretrained baseline.

G. Harshitha, Vasudeva, Nisha P. Poojary et al. · 0 citations
Preprint Aug 2026

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

This paper proposes a general optimization framework that combines a vocabulary pruning method with a targeted fine-tuning protocol for MNMT models, and reduces the vocabulary size from over 128,000 to approximately 10,000 tokens, enabling a 60% memory saving without any loss in performance.

A. A. Aliane, N. Semmar, H. Aliane · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.