Jul 2026· Signal, Image and Video Processing· Vol 20· 0 citations· 30 references
Computer Science
TL;DR
MoCoMER-V1 is presented, an attention-guided self-supervised framework to improve representation learning through addition of channel and spatial attention to a Momentum Contrast (MoCo) pipeline, which achieves competitive performance compared to self-supervised HMER and even surpasses several fully supervised baselines.
Handwritten character recognition plays a crucial role in optical character recognition systems, particularly for low-resource and structurally complex scripts such as Tamil. Despite significant advances in deep learning, accurate recognition of handwritten Tamil characters remains challenging due to large variations i...
K. Manoj, M. Iyapparaja· Intelligent Data Analysis· 0 citations
Handwritten text recognition resources are often small and distributed across collections that differ in language, script, document structure, and annotation format, making joint page-level training difficult. We propose ExpertHTR, a unified vision-language framework that addresses this problem through complementary su...
Nam Hoai Dang, H. Nguyen, Quang Huu Hieu et al.· 0 citations
This paper proposes a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism, which trains a Transformer encoder with multi-head self-attention that models long-range context far more effectively than bidirectional recurrent layers.
Lalita Kumari· International journal on eme...· 0 citations
An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...
Bilal Abdulrahman, Farhan Mohamed· Journal of Human Centered Te...· 0 citations
This is the first study to introduce both a real handwritten Urdu word dataset and a diffusion-generated synthetic dataset, and develops a unified word recognition model trained jointly on handwritten and printed Urdu word data, leading to improved recognition robustness and performance.
Wahid Hussain, Shahbaz Hassan, I. Hassan et al.· IEEE Access· 0 citations
Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...
Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.