Skip to content
Open access

Sequence Labeling in Urdu Social Media Texts: Data Annotation and Transformer-Based Deep Learning Models

Aug 2026 · ACM Transactions on Asian and Low-Resource Language Information Processing · 0 citations · 17 references

TL;DR

A transformer-based hybrid architecture is proposed, XLM-R–CNN–BiLSTM–CRF, which integrates XLM-R embeddings for contextualized representation, convolutional neural networks (CNNs) for local feature extraction, bidirectional long short-term memory (BiLSTM) networks for sequential modeling, and a CRF layer for optimal sequence prediction.

Abstract

Sequence labeling tasks such as part-of-speech (POS) tagging and named entity recognition (NER) have advanced significantly in high-resource languages and well-structured texts. However, low-resource languages, such as Urdu, face unique challenges due to limited annotated resources, complex morphology, and the informal, noisy nature of social media text. To address this, we introduce two relatively large-scale annotated datasets of Urdu tweets: a POS tagging dataset with 39 syntactic tags and an NER dataset covering three major entity types (person, location, and organization), capturing the linguistic diversity of user-generated content. We benchmark these datasets using a spectrum of approaches, from traditional conditional random fields (CRFs) to deep neural architectures with static embeddings and fine-tuned pre-trained language models (PLMs). Building on these findings, we propose a transformer-based hybrid architecture, XLM-R–CNN–BiLSTM–CRF, which integrates XLM-R embeddings for contextualized representation, convolutional neural networks (CNNs) for local feature extraction, bidirectional long short-term memory (BiLSTM) networks for sequential modeling, and a CRF layer for optimal sequence prediction. Our approach achieves state-of-the-art performance, with F1 scores of 95.39% for POS tagging and 91.91% for NER, significantly surpassing strong baselines. These resources and methods advance sequence labeling for Urdu while providing insights for other low-resource, noisy languages.

Read PDF

Similar papers

Open access Sep 2026

BENCHMARKING DEEP LEARNING MODELS FOR FEW-SHOT CLASSIFICATION OF KAZAKH NEWS TEXTS UNDER LIMITED ANNOTATION

Deep learning methods and transformer architectures have fundamentally reshaped the field of natural language processing; however, these advancements remain unevenly distributed. While high-resource languages like English benefit from large-scale benchmarks, robust text classification for low-resource languages, such a...

D. Marlambekov, Асель Сейтказиевна Ахметова, S. Torekul et al. · 0 citations
Open access 2026

Attention-Augmented Hybrid CNN- BiLSTM Architecture for Real-Time Sentiment Classification of Twitter Text

This study proposes a hybrid deep-learning framework that combines convolutional feature extraction with a bidirectional recurrent encoder and an attention mechanism to categories tweets as positive, neutral, or negative, and indicates that the learned representations generalize reasonably well across domains rather th...

K. Vadivelan, M. Rajan · 0 citations
Review Aug 2026

A Review of Deep Learning-Based Text Classification Research

The exponential growth of textual data on social media and information networks poses a significant challenge to extracting valuable information. Text classification, a core task in Natural Language Processing (NLP), is essential for organizing and categorizing such data. Deep learning has emerged as an effective a...

Ran Jin, Ya Wang, Tianzi Wu et al. · 0 citations
Review Open access Sep 2026

A structured and comprehensive review of Sino-Tibetan languages PoS taggers.

This study examines existing studies on PoS tagging of several Sino-Tibetan languages, explains the datasets and models used, and compares their stated performance, and discusses the limits of the prior approaches.

Bedawati Basumatary, Shikhar Kumar Sarma, Kuwali Talukdar et al. · 0 citations
Open access Sep 2026

Sentiment classification of stock forum text based on a parameter-decoupled ERNIE-Transformer architecture

To address the challenges of diverse domain-specific terminology, highly colloquial expressions, and limited annotated samples in sentiment analysis of stock forum texts, this study proposes an ERNIE-Transformer sentiment classification model that integrates ERNIE and Transformer architectures. First, a systematic data...

Xiu-Mei Li, Fei Chen, Wen-Chao Ling et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.