Skip to content
Open access

Enhancing Readability of Telugu Text Summarization Using Multi-scale Attention and Bio-inspired Optimization

Jul 2026 · ACM Transactions on Asian and Low-Resource Language Information Processing · Vol 25, pp. 1-31 · 0 citations · 22 references

TL;DR

The proposed MLOA-MA-ASeqNet architecture, a Multi-scale Attention and Adaptive Sequence-to-Sequence Network whose hierarchical encoder operates simultaneously at word, phrase and sentence level granularity, achieves the highest average score across fluency, adequacy, coherence and readability.

Abstract

Telugu ranks among the most widely spoken language in South Asia, yet it remains conspicuously underrepresented in automatic summarization research. This neglect is not arbitrary; it reflects three genuine problems viz., Telugu’s agglutinative nature, syntactic representation, and lack of annotated corpora. To address these problems within a unified framework, this paper introduces MLOA-MA-ASeqNet architecture. The architectural core is MA-ASeqNet, a Multi-scale Attention and Adaptive Sequence-to-Sequence Network whose hierarchical encoder operates simultaneously at word, phrase and sentence level granularity. The optimization component, MLOA, is a Modified Lyrebird Optimization Algorithm that replaces manually configured, English-centric hyperparameter defaults with a principled population-based search. Experiments across three datasets – Telugu News NLP, Telugu Books, and TeSum; show consistent and statistically improvements over Seq2Seq, Transformer, T5 and Gemma baselines on ROUGE scores (ROUGE-L of 0.53). Qualitative analysis of the generated summaries was performed through a blind evaluation by five native Telugu speakers from diverse professional backgrounds. The summaries were evaluated on four parameters: fluency, adequacy, coherence, and readability. The proposed MLOA-MA-ASeqNet achieved the highest average score of 4.60 across fluency, adequacy, coherence and readability; surpassing all four baselines on every dimension. All pairwise differences were statistically significant according to the Wilcoxon signed-rank test with Bonferroni correction (p < 0.05 in all cases).

Read PDF

Similar papers

Conference Jul 2026

Bharat Sum: OCR-Enabled Multilingual News Summarization and Bias Analysis Framework

India is home to an incredible number of languages which leads to the production of substantial amounts of news articles provided in their regional languages including Telugu, Tamil, Hindi, Bengali, Kannada, and Malayalam. Unfortunately, current systems for processing this data do so independently, rather than as part of a comprehensive framework utilizing all the necessary components in one queue of processing pipelines; OCR extraction, Translation, Summarization, and Bias Detection must all be completed one at a time, and do not allow for seamless data passing between functions. In this paper, we will present Bharat Sum, a multilingual news summarization and bias detection system that incorporates OCR capabilities into five different types of processing stages; OCR Text Extraction using Tesseract, Automatic Language Detection using LangDetect, Topic Segmented Abstractive Summarisation using mT5, Translation to English using mBART, and Sentiment based Bias Classification using DistilBERT - all accessible through a single scalable architecture running on commodity hardware and implemented via Streamlit. Our testing involved 150 news articles covering each of the five languages named above. The results achieved were as follows; OCR extraction accuracy of 89.7%, Language Detection Accuracy of 94.2%, Summarisation Quality (ROUGE-L F1) of 92.1% Translation consistency of 91.4% and Sentiment Classification Accuracy of 88.6%. The average end-to-end processing time was between 10 and 16 seconds. Our analysis of Bharat Sum has revealed that it significantly outperforms previous single function systems by providing an Integrated, Real-Time Multilingual Processing capability which currently does not exist in this context. Bharat Sum has the potential to address significant gaps in the research literature regarding the analysis of Integrated Multilingual Media, and will likely serve as an economically viable solution for organisations conducting Digital Journalism, Media Monitoring, or Accessing Multilingual Information.

Farooq Sunar Mohammad, E.Sneha, B.Kavya et al. · 0 citations
Book Open access Jul 2026

Graph-Enhanced Sentence Retrieval for Multi-Document Summarization in Low-Resource Languages

Multi-document summarization for low-resource languages faces a critical trade-off: large language models are computationally prohibitive for most institutions, while smaller models suffer from severe hallucination in abstractive generation. We address this through extractive sentence retrieval, which guarantees faithfulness while operating within constrained computational budgets. Our approach combines language-adaptive mixture-of-experts embeddings with graph neural networks that model discourse structure, addressing linguistic challenges across typologically diverse low-resource languages. With only 3.2M trainable parameters, our model requires 28 times less training time than comparable transformer-based approaches, making it practical for single-GPU environments typical in resource-limited settings. We demonstrate applicability to some main Southeast Asia (SEA) countries including Vietnam, Thailand, Laos, Indonesia, and Malaysia, representing three distinct language families: Austroasiatic, Kra-Dai, and Austronesian.

Xuan-Hung Le, Thi Toan Do, Hoang-Quynh Le · 0 citations
Conference Aug 2026

Data-Centric Evaluation of Arabic Abstractive Summarization Using a Large-scale Curated News Corpus

This paper introduces MAAD, a high-quality, carefully constructed and curated by the authors large-scale Arabic dataset for abstractive news summarisation. The authors selected a high-quality subset of 50,000 articles from the dataset Original, which contains 602,792 articles. To maintain the quality, diversity, and training suitability of the subset, the subset underwent a multi-stage preprocessing pipeline involving noise removal, duplicate filtering, linguistic normalisation, and expert validation. The experimental evaluation was executed in two phases. In the first phase, three transformer-based models (ArabicT5, AraBART, and mT5) were evaluated on a controlled subset of 1,110 articles to establish fair baseline comparisons among models, where ArabicT5 achieved the best performance (ROUGE-1: 23.64, ROUGE-2: 11.82, ROUGE-L: 22.10). In the second phase, ArabicT5-base was trained on all 50,000 articles to evaluate scalability, achieving substantially improved results of 68.4, 52.3, and 64.1, respectively, with a BLEU score of 58.7. The findings emphasise the significance of scale, effective preprocessing, and the benefits of Arabic-specific pretraining on the quality of summarisation. Moreover, a human evaluation on 500 randomly sampled instances verified fluency and adequacy scores of 4.86 and 4.35, respectively, with a strong inter-annotator agreement (Cohen's Kappa: 0.78 and 0.74). Overall, the findings indicate that MAAD is a reliable and scalable dataset with strong potential to serve as a benchmark for Arabic abstractive summarisation and to support the development of robust transformer-based models.

M. Al-Nahari, Ayedh abdulaziz Mohsen, Nada Abdu Al-Humidi et al. · 0 citations
Open access Jul 2026

LSTM-Based Classification of Indonesian Regional Song Lyrics by Language

This study successfully proposes a Long Short-Term Memory (LSTM)-based model for automatic classification of Indonesian regional song lyrics by language, demonstrating that LSTM effectively captures sequential linguistic patterns and contextual relationships within regional languages.

Muhammad Rizky, Anandita Priatama, Aviv Yuniar Rahman et al. · 0 citations
Open access Jul 2026

Deep learning-based extractive and abstractive summarization for the Azerbaijani language

This article investigates extractive and abstractive text summarization for the Azerbaijani language, a low-resource and underrepresented language in natural language processing. While the underlying modeling approaches are well established, their application to Azerbaijani summarization remains largely unexplored due to the scarcity of large-scale datasets and prior empirical studies. To address this gap, we conduct a systematic evaluation of both extractive methods based on sentence ranking and an abstractive approach using a fine-tuned mT5-base model. Our experiments are carried out on a large-scale dataset comprising over 115,000 Azerbaijani news articles paired with human-written summaries. The models are evaluated using standard automatic metrics, including Recall-Oriented Understudy for Gisting Evaluation (ROUGE), Bilingual Evaluation Understudy (BLEU), and Metric for Evaluation of Translation with Explicit ORdering (METEOR), yielding strong results that highlight the benefits of task specific fine-tuning for abstractive summarization, while also demonstrating the competitiveness of extractive baselines. In addition, we analyze the impact of long input sequences and discuss architectural and dataset-related limitations affecting performance. Overall, this study provides a comprehensive empirical baseline for Azerbaijani text summarization and serves as a reference point for future research in low-resource summarization and related Azerbaijani Natural Language Processing (NLP) applications.

Mir Amir Pashayev, S. Rustamov · 0 citations
Book Jul 2026

NERBench-Chhattisgarh: A Multi-Family NER Dataset for Low-Resource Indic Languages

We present NERBench-Chhattisgarh, a gold-standard Named Entity Recognition (NER) dataset covering seven under-resourced languages spoken in Central India: Baigani, Chhattisgarhi, Surgujia, Sadri, Kudukh, Halbi, and Gondi. Addressing the digital divide for tribal languages, our corpus comprises 166,444 annotated tokens across 8,391 sentences, spanning both Indo-Aryan and Dravidian language families. The dataset features high lexical sparsity and a "nature-centric" ontology of 22 entity types based on the CLIA Phase-II schema, capturing culturally specific entities often absent in standard benchmarks. To establish a benchmark for language variety-aware information access, we evaluate three multilingual encoders: mBERT, XLM-RoBERTa, and IndicBERT, using two adaptation strategies: direct parameter-efficient fine-tuning (LoRA) and a Chhattisgarhi-Pivot Adaptive Pre-training (CPAP) approach. Our results show a morphological barrier: while pivot-based adaptation facilitates transfer for distant languages through script alignment, it induces negative transfer in morphologically complex agglutinative languages like Gondi. We publicly release the dataset, code, and adapted model checkpoints. https://github.com/Rajesh-NLP/NER-Chhattisgarh to support future research in inclusive Information Retrieval.

R. Mundotiya · 0 citations