Skip to content
Open access

A Hybrid Vision Mamba and Transformer Architecture for Offline Recognition of Handwritten Marathi Characters

Jul 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 44 references

TL;DR

A Hybrid Vision Mamba and Transformer (HVMT), framework for stronger offline handwritten Marathi character recognition and a good fit for things like intelligent document analysis, handwritten document digitization, archival preservation, and several other Indic script recognition tasks and more.

Abstract

Offline handwritten Marathi character recognition is still kind of hard research problem because there is so much variability within the same class ,and between classes they can look a bit similar ,also the strokes are complex and different people write in their own style. A lot of CNN and Transformer like methods either do not really capture long range relationships well enough, or they end up being too heavy computationally, you know not so efficient. So in this paper we suggest a Hybrid Vision Mamba and Transformer (HVMT), framework for stronger offline handwritten Marathi character recognition. The HVMT idea combines the hierarchical feature extraction power of Vision Mamba, which uses selective state-space modeling, with the contextual representation learning of a smaller Transformer encoder, and inside that encoder we use Multi-Head Self-Attention. Experiments are done on the public MHCD_GIETV2 dataset, where handwritten Marathi characters are collected from writers in different age groups and with diverse writing styles. Before training the images are turned into grayscale, then normalized, resized, and also augmented, to help the model generalize better. The proposed HVMT is compared with CNN, ResNet-50, EfficientNet-B0, ConvNeXt-Tiny, Vision Transformer (ViT-B/16), Swin Transformer-Tiny, and Vision Mamba, all under the same experimental setup. Experimental results show that the proposed framework achieved accuracy 87.11% , precision 87.09% , recall 87.10% and F1-score 87.09% which is better than the compared architectures. At the same time it only uses 26.9 million parameters, 2.6 GFLOPs, and inference time 1.305 ms per image. In other words, the HVMT framework seems to strike a workable tradeoff between recognition precision and compute efficiency. Because of this it is a good fit for things like intelligent document analysis, handwritten document digitization , archival preservation , and several other Indic script recognition tasks and more.

Read PDF

Similar papers

Open access 2026

Multi-Scale Transformer-Based Lexicon-Guided Handwritten Text Recognition Using Adaptive Feature Fusion

This paper proposes a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism, which trains a Transformer encoder with multi-head self-attention that models long-range context far more effectively than bidirectional recurrent layers.

Lalita Kumari · 0 citations
Open access Aug 2026

Transformer Framework with Dynamic Learning Rate Optimisation for Offline Devanagari Character Recognition

A unified multi-attention framework was developed that explicitly optimises structural handwritten character script recognition through an adaptive multi-phase learning-rate schedule, incorporating global and hierarchical local transformer operations (ViT and Swin Transformer), respectively.

P. M. Kakde, S. Gulhane · 0 citations

Offline Bengali Writer Verification by PDF-CNN and Siamese Net

This paper deals with offline writer verification on complex handwriting patterns, where handcrafted feature PDFs are hybridized with auto-derived CNN features and fed into a Siamese neural network for writer verification.

Chandranath Adak, SimoneMarinai, BidyutB.Chaudhuri et al. · 0 citations
Open access Jul 2026

HMAFNet: A Hierarchical Multi-Scale Attention Fusion Network for Offline Handwritten Odia Compound Character Recognition

Handwriting based offline Odia compound character recognition could be challenging mainly due to the complex structure of the conjunct characters along with having significant variations in handwriting style. Moreover, touching characters, discontinuous strokes, degradation in documents, and so forth make the problem e...

Sachikanta Dash · 0 citations
Open access 2026

Diffusion-Enhanced NAT–BART Vision Language Transformer for Unified Urdu Word Recognition

This is the first study to introduce both a real handwritten Urdu word dataset and a diffusion-generated synthetic dataset, and develops a unified word recognition model trained jointly on handwritten and printed Urdu word data, leading to improved recognition robustness and performance.

Wahid Hussain, Shahbaz Hassan, I. Hassan et al. · 0 citations
Open access Aug 2026

Character-Based Arabic Offline Handwritten Text Recognition Using Faster R-CNN

These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings.

Sofiane Medjram, Ruwaidah Saud Alnejaidi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.