Skip to content

Mocomer-v1: attention-guided contrastive pretraining for robust handwritten mathematical expression recognition

Jul 2026 · Signal, Image and Video Processing · Vol 20 · 0 citations · 30 references
Computer Science

TL;DR

MoCoMER-V1 is presented, an attention-guided self-supervised framework to improve representation learning through addition of channel and spatial attention to a Momentum Contrast (MoCo) pipeline, which achieves competitive performance compared to self-supervised HMER and even surpasses several fully supervised baselines.

View source

Similar papers

Aug 2026

TamilLite-Gan: A lightweight graph attention network for handwritten Tamil character recognition using enhanced single shot optimization

Handwritten character recognition plays a crucial role in optical character recognition systems, particularly for low-resource and structurally complex scripts such as Tamil. Despite significant advances in deep learning, accurate recognition of handwritten Tamil characters remains challenging due to large variations i...

K. Manoj, M. Iyapparaja · 0 citations
#machine learning Preprint Sep 2026

ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts

Handwritten text recognition resources are often small and distributed across collections that differ in language, script, document structure, and annotation format, making joint page-level training difficult. We propose ExpertHTR, a unified vision-language framework that addresses this problem through complementary su...

Nam Hoai Dang, H. Nguyen, Quang Huu Hieu et al. · 0 citations
Open access 2026

Multi-Scale Transformer-Based Lexicon-Guided Handwritten Text Recognition Using Adaptive Feature Fusion

This paper proposes a Multi-Scale Transformer-Based Lexicon-Guided HTR framework built around an Adaptive Feature Fusion (AFF) mechanism, which trains a Transformer encoder with multi-head self-attention that models long-range context far more effectively than bidirectional recurrent layers.

Lalita Kumari · 0 citations
Open access Aug 2026

Improving Right to Left Cursive Handwritten Text Recognition in Historical Manuscripts Using Learnable Edge Features and Channel Attention

An edge-aware line-level HTR framework that extends a CNN-Transformer baseline with a learnable edge-extraction channel and Squeeze-and-Excitation channel attention and shows that combining learnable structural cues with channel-wise attention has improved robustness for degradation-prone historical manuscript collecti...

Bilal Abdulrahman, Farhan Mohamed · 0 citations
Open access 2026

Diffusion-Enhanced NAT–BART Vision Language Transformer for Unified Urdu Word Recognition

This is the first study to introduce both a real handwritten Urdu word dataset and a diffusion-generated synthetic dataset, and develops a unified word recognition model trained jointly on handwritten and printed Urdu word data, leading to improved recognition robustness and performance.

Wahid Hussain, Shahbaz Hassan, I. Hassan et al. · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.