Skip to content
Open access

Multi-Scale Transformer-Based Lexicon-Guided Handwritten Text Recognition Using Adaptive Feature Fusion

2026 · International journal on emerging technologies · 0 citations

TL;DR

This paper proposes a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism, which trains a Transformer encoder with multi-head self-attention that models long-range context far more effectively than bidirectional recurrent layers.

Abstract

Offline handwritten text recognition (HTR) remains one of the most challenging problems in document image analysis because of cursive writing, character overlap, unconstrained writing styles, long inter-character dependencies and background degradation. Classical CNN-BLSTM pipelines and single-scale attention models extract features at a fixed resolution and therefore struggle to jointly represent thin strokes of small characters and the global shape of large or connected characters. In this paper we propose a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism. A multi-scale convolutional backbone extracts shallow, middle and deep feature maps using parallel 3×3, 5×5 and 7×7 receptive fields. The AFF module learns content-dependent soft weights and channel attention to fuse these heterogeneous features into a single scale-balanced representation, replacing the fixed feature map used by previous lexicon-and-attention systems. A Transformer encoder with multi-head self-attention then models long-range context far more effectively than bidirectional recurrent layers. The network is trained end-to-end with the Connectionist Temporal Classification (CTC) objective and decoded with a lexicon-guided beam search that combines an n-gram prior with edit-distance dictionary matching. On the IAM line-level benchmark the proposed model attains a character error rate (CER) of 3.08% and a word error rate (WER) of 7.24%, improving over our lexicon-and-attention baseline (4.15% CER / 9.72% WER) by 25.8% and 25.5% relative, respectively. On RIMES we obtain 2.71% CER / 7.92% WER, and on the George Washington collection 5.82% CER / 13.05% WER, again outperforming the baseline.

Read PDF

Similar papers

Open access Aug 2026

Improving Right to Left Cursive Handwritten Text Recognition in Historical Manuscripts Using Learnable Edge Features and Channel Attention

An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...

Bilal Abdulrahman, Farhan Mohamed · 0 citations
Open access Jul 2026

HMAFNet: A Hierarchical Multi-Scale Attention Fusion Network for Offline Handwritten Odia Compound Character Recognition

Handwriting based offline Odia compound character recognition could be challenging mainly due to the complex structure of the conjunct characters along with having significant variations in handwriting style. Moreover, touching characters, discontinuous strokes, degradation in documents, and so forth make the problem e...

Sachikanta Dash · 0 citations
Open access Jul 2026

A Hybrid Vision Mamba and Transformer Architecture for Offline Recognition of Handwritten Marathi Characters

A Hybrid Vision Mamba and Transformer (HVMT), framework for stronger offline handwritten Marathi character recognition and a good fit for things like intelligent document analysis, handwritten document digitization, archival preservation, and several other Indic script recognition tasks and more.

S. Khandakhani, Sachikanta Dash, Sasmita Padhy et al. · 0 citations
Jul 2026

Mocomer-v1: attention-guided contrastive pretraining for robust handwritten mathematical expression recognition

MoCoMER-V1 is presented, an attention-guided self-supervised framework to improve representation learning through addition of channel and spatial attention to a Momentum Contrast (MoCo) pipeline, which achieves competitive performance compared to self-supervised HMER and even surpasses several fully supervised baseline...

Sandip Pramanik, Nibaran Das · 0 citations
Open access 2026

Diffusion-Enhanced NAT–BART Vision Language Transformer for Unified Urdu Word Recognition

This is the first study to introduce both a real handwritten Urdu word dataset and a diffusion-generated synthetic dataset, and develops a unified word recognition model trained jointly on handwritten and printed Urdu word data, leading to improved recognition robustness and performance.

Wahid Hussain, Shahbaz Hassan, I. Hassan et al. · 0 citations
Aug 2026

TamilLite-Gan: A lightweight graph attention network for handwritten Tamil character recognition using enhanced single shot optimization

Handwritten character recognition plays a crucial role in optical character recognition systems, particularly for low-resource and structurally complex scripts such as Tamil. Despite significant advances in deep learning, accurate recognition of handwritten Tamil characters remains challenging due to large variations i...

K. Manoj, M. Iyapparaja · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.