Aug 2026· ACM Transactions on Asian and Low-Resource Language Information Processing· 0 citations· 17 references
TL;DR
A transformer-based hybrid architecture is proposed, XLM-R–CNN–BiLSTM–CRF, which integrates XLM-R embeddings for contextualized representation, convolutional neural networks (CNNs) for local feature extraction, bidirectional long short-term memory (BiLSTM) networks for sequential modeling, and a CRF layer for optimal sequence prediction.
Abstract
Sequence labeling tasks such as part-of-speech (POS) tagging and named entity recognition (NER) have advanced significantly in high-resource languages and well-structured texts. However, low-resource languages, such as Urdu, face unique challenges due to limited annotated resources, complex morphology, and the informal, noisy nature of social media text. To address this, we introduce two relatively large-scale annotated datasets of Urdu tweets: a POS tagging dataset with 39 syntactic tags and an NER dataset covering three major entity types (person, location, and organization), capturing the linguistic diversity of user-generated content. We benchmark these datasets using a spectrum of approaches, from traditional conditional random fields (CRFs) to deep neural architectures with static embeddings and fine-tuned pre-trained language models (PLMs). Building on these findings, we propose a transformer-based hybrid architecture, XLM-R–CNN–BiLSTM–CRF, which integrates XLM-R embeddings for contextualized representation, convolutional neural networks (CNNs) for local feature extraction, bidirectional long short-term memory (BiLSTM) networks for sequential modeling, and a CRF layer for optimal sequence prediction. Our approach achieves state-of-the-art performance, with F1 scores of 95.39% for POS tagging and 91.91% for NER, significantly surpassing strong baselines. These resources and methods advance sequence labeling for Urdu while providing insights for other low-resource, noisy languages.
Deep learning methods and transformer architectures have fundamentally reshaped the field of natural language processing; however, these advancements remain unevenly distributed. While high-resource languages like English benefit from large-scale benchmarks, robust text classification for low-resource languages, such a...
D. Marlambekov, Асель Сейтказиевна Ахметова, S. Torekul et al.· Herald of Kazakh-British tec...· 0 citations
This study proposes a hybrid deep-learning framework that combines convolutional feature extraction with a bidirectional recurrent encoder and an attention mechanism to categories tweets as positive, neutral, or negative, and indicates that the learned representations generalize reasonably well across domains rather th...
K. Vadivelan, M. Rajan· International journal of res...· 0 citations
The exponential growth of textual data on social media and information networks poses a significant challenge to extracting valuable information. Text classification, a core task in Natural Language Processing (NLP), is essential for organizing and categorizing such data. Deep learning has emerged as an effective a...
Ran Jin, Ya Wang, Tianzi Wu et al.· Recent Advances in Computer...· 0 citations
This study examines existing studies on PoS tagging of several Sino-Tibetan languages, explains the datasets and models used, and compares their stated performance, and discusses the limits of the prior approaches.
Bedawati Basumatary, Shikhar Kumar Sarma, Kuwali Talukdar et al.· Frontiers in Artificial Inte...· 0 citations
To address the challenges of diverse domain-specific terminology, highly colloquial expressions, and limited annotated samples in sentiment analysis of stock forum texts, this study proposes an ERNIE-Transformer sentiment classification model that integrates ERNIE and Transformer architectures. First, a systematic data...
Xiu-Mei Li, Fei Chen, Wen-Chao Ling et al.· Journal of Electrical System...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.