Aug 2026· International Conference on Machine Vision and Deep Learning· Vol 14326, pp. 143262U - 143262U-8· 0 citations· 5 references
Engineering
TL;DR
This paper addresses the issues of unstable recognition, semantic misalignment, and excessive inference costs that often occur in "OCR + machine translation" under complex backgrounds, vertical-horizontal mixed layouts, and low-resolution conditions by constructing a set of automated translation optimization framework with clear engineering mechanisms.
Abstract
In response to the rapidly growing demand for Japanese visual text translation in scenarios such as cross-border e-commerce, games, and online education, this paper addresses the issues of unstable recognition, semantic misalignment, and excessive inference costs that often occur in "OCR + machine translation" under complex backgrounds, vertical-horizontal mixed layouts, and low-resolution conditions. It has constructed a set of automated translation optimization framework with clear engineering mechanisms. Starting from the complexity index of the structure, it introduces a character recognition loss that combines structural perception and uncertainty weighting, and explicitly models the cascading distortion of visual uncertainty on translation quality in multiple scenarios. At the visual-language decoding end, it designs character/line/block three-level alignment and consistency regularization, and adds an ambiguity-driven gating mechanism for context. At the system level, it achieves multi-configured inference under efficiency constraints through "difficulty perception + budget-driven" dynamic routing. Experiments show that the complete system improves BLEU from 32.5 to 36.8, and the semantic consistency score increases from 78.0 to 85.6. The lightweight configuration reduces FLOPs from 45.3 G to 18.7 G while compressing the end-to-end latency from 185 ms to 92 ms, demonstrating a better synergy of accuracy and efficiency. This provides a feasible path for the engineering deployment of Japanese visual text in cross-terminal and multi-scenario conditions.
Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...
Y.-L. Diao, W. Gao· Advanced Electromagnetics· 0 citations
To address the issues of amplified recognition errors and input distribution mismatch in Japanese speech translation scenarios, a "translation-aware ASR" collaborative optimization method was constructed. During the speech recognition training stage, a joint loss with translation target constraints was introduced, and...
Xiao-Yi Hou· International Conference on...· 0 citations
Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition,...
Guang-Zhan Huang, Yong-Shuo Zhang, Bing-Tao Fu et al.· 0 citations
Artistic Text Recognition (ATR) remains challenging because word images often combine decorative fonts, curved layouts, object-like characters, clutter, and severe distortions. This paper studies WordArt-V1.5 as a standardized benchmark for this setting and evaluates recent scene and artistic text recognizers under a c...
L. A. Dias, Henrique A. Schulz, Rafael Tadeu Machado de Miranda et al.· 0 citations
SPIRAL (Self-improving Path Integration and Realignment), a self-supervised alignment framework that closes this gap using only the model's own text-path behavior as supervision, requiring no external teachers or additional annotations, generalize to out-of-domain benchmarks, confirming that effective VTC hinges on ali...
Tianyu Liang, Xiangxi Zheng, Yilin Wang et al.· 2 citations
This paper addresses the task of automatic identification of German word forms and constructs an end-to-end model framework based on deep neural networks that employs character-level and subword-level dual-channel feature representations, and combines encoder-decoder architecture, scaled dot-product attention, and posi...
Bo Wang· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.