Leveraging Multilingual Pre-Trained Transformers for Abstractive Summarization of Malayalam Educational Text: A Study on Social-Sum-Mal
Abstract
Malayalam, a classical Dravidian language spoken by approximately 38 million people in the Indian state of Kerala, remains severely underrepresented in natural language processing research. This paper presents the systematic study of abstractive text summarization for Malayalam educational documents using mT5-base, a massively multilingual sequence-to-sequence transformer model. The present study employs Social-Sum-Mal, a native Malayalam dataset comprising 2,000 educational text-summary pairs derived from social studies textbooks. Three summary types available in the dataset — long summary, extreme summary, and answer summary — are analyzed separately to determine their suitability for training abstractive summarization models. The model is fine-tuned for 15 epochs with a learning rate of 1×10–4, batch size of 2 with gradient accumulation of 16 steps, on a Tesla T4 GPU. Evaluation on the test set yields ROUGE-1 of 9.81, ROUGE-2 of 2.75, ROUGE-L of 9.81, and BERT Score F1 of 88.18. Analysis reveals that long summary produces the highest inter-summary ROUGE agreement (ROUGE-1: 36.00) among the three summary types, making it the most suitable training target. The high BERT Score of 88.18 confirms that the model generates semantically accurate Malayalam summaries despite the limited training data, establishing a strong baseline for Malayalam abstractive summarization research.