Natural language processing optimizes Japanese speech recognition technology for translation applications
Abstract
To address the issues of amplified recognition errors and input distribution mismatch in Japanese speech translation scenarios, a "translation-aware ASR" collaborative optimization method was constructed. During the speech recognition training stage, a joint loss with translation target constraints was introduced, and in the inference stage, translation consistency constraints were used to achieve joint decision-making between ASR and MT. At the same time, semantic perception modeling and cross-modal intermediate representation sharing mechanisms were utilized to enhance the adaptability of the recognition results to translation decoding. Experimental results show that, under the condition of maintaining controllable computational overhead, this method reduces CER by approximately 1.5%–2.0% compared to the baseline model, and the translation BLEU score increases by more than 2.5, demonstrating the practical value of this method in the application of Japanese speech translation engineering.