Skip to content
Open access

Abstractive Summarization of Long Educational Transcripts in the Era of GenAI

Aug 2026 · Journal of Computer Science · 0 citations · 56 references

Abstract

: The surge in online video data presents the imperative requirement of summarizing it for information compression, efficient understanding, and to support decision-making. Recent research studies have predominantly utilized structured input data for summarization. This paper explores the current state-of-the-art models in summarizing unstructured, conversational, and long transcript data. We further apply a transfer learning approach using a fine-tuned BART model, pre-trained on the SAMSum dataset. Hyperparameter tuning and selective layer freezing are applied to optimise model performance. This study focuses on abstractive summarization of long video transcript data. Our methodology integrates ChatGPT to generate reference abstractive summaries from long transcript data. A semi-automated pipeline using TextRank is proposed for reference summary generation. The proposed fine-tuned model shows a significant increase in ROUGE scores over the baseline model. The findings suggest that the transfer learning approach is effective for abstractive summarization of long transcript data in real-world conversational domains.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.