Skip to content
Review Open access

Stemming techniques for resource-poor languages: a review of methods, challenges, and applications

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 68 references

TL;DR

A systematic review of stemming techniques across Afro-Asiatic, Indo-Aryan, Turkic, and Uralic language families, with explicit acknowledgment of coverage limitation, points out the strengths and weaknesses of current approaches and provides insights into potential areas for future investigation and developments in language-sensitive stemming systems.

Abstract

The process of stemming is a fundamental part of pre-processing in NLP systems in resource-poor and morphologically complex languages in which linguistic data and annotated corpora are limited. This survey presents a a systematic review of stemming techniques across Afro-Asiatic, Indo-Aryan, Turkic, and Uralic language families, with explicit acknowledgment of coverage limitation and it covers rule-based, statistical, unsupervised, hybrid, and emerging neural-assisted approaches. The evolution of stemming is analyzed in the broader context of NLP paradigms, from early suffix-stripping algorithms to modern sub-word-aware and transformer-influenced models. Special emphasis has been given on resource-poor languages such as Arabic, Urdu, Bengali, Assamese, and other low-resource Indo-Aryan, Afro-Asiatic, and agglutinative languages. For these languages performing explicit stemming remains a critical task despite advances in deep learning. This study further reviews application-driven impacts of stemming in information retrieval, text classification, sentiment analysis, topic modeling, and semantic similarity estimation. A unified comparison framework is presented which has incorporated benchmark datasets, shared-task resources, and multi-dimensional evaluation metrics, including precision, recall, F-measure, MAP, NMI, and stemming-specific error indices. Using careful comparative evaluation, this study points out the strengths and weaknesses of current approaches and provides insights into potential areas for future investigation and developments in language-sensitive stemming systems.

Read PDF

Similar papers

Review Open access Sep 2026

A structured and comprehensive review of Sino-Tibetan languages PoS taggers

Parts of Speech (PoS) tagging is a fundamental activity in Natural Language Processing (NLP). It consists in assigning an appropriate grammatical category to each word in a text, such as noun, verb, adjective or adverb. It serves as a crucial preprocessing step in many NLP applications such as machine translation, info...

Bedawati Basumatary, Shikhar Kumar Sarma, Kuwali Talukdar et al. · 0 citations
Review Open access Aug 2026

Applications of Natural Language Processing: A Comprehensive Study

A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.

P. Kalaiselvi · 0 citations
2026

Efficient Adaptation of English Language Models for Morphologically Rich and Underrepresented Languages: The Case of Arabic

A resource-efficient adaptation of the English-pretrained ModernBERT for Arabic, employing continued pretraining on large Arabic corpora followed by lightweight head-only fine-tuning with a frozen encoder, demonstrating that modern English encoder architectures can be efficiently transferred to Arabic through language-...

Ahmed Samy Eldamaty, M. Abdelrahman, Mohamed Mostafa Ibrahim Elbehery et al. · 0 citations
2026

MaitH 1.0: A Parallel Corpus and Baseline for Low-Resource Maithili-Hindi Translation

A corpus containing both manually curated and synthetically generated sentences for low-resource Indian languages, such as Maithili is contributed and it is demonstrated that, even with a smaller corpus size, high-quality, task-specific data significantly enhance translation accuracy for low-resource Indian languages,...

Kamanksha Prasad Dubey, C. Maurya, Kumar Padmanabh · 0 citations
Open access Aug 2026

A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo

The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.

Maureen Otieno, L. Wanzare, Calvins Otieno · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.