Aug 2026· Athens Journal of Τechnology & Engineering· Vol 13, pp. 211-228· 1 citation
TL;DR
A Generative Pre-trained Transformers-only model that relies on prompt engineering against a Retrieval-Augmented Generation model that incorporates external university documents, specifically program flyers and a module handbook, integrated using Langchain are evaluated.
Abstract
While Large Language Models (LLMs) have demonstrated impressive capabilities in general natural language processing, their accuracy often diminishes in domain-specific contexts where precise, factual responses are crucial. This study addresses this limitation within the higher education sector by comparing two approaches to handling university-specific queries. We evaluate a Generative Pre-trained Transformers (GPT)-only model that relies on prompt engineering against a Retrieval-Augmented Generation (RAG) model that incorporates external university documents, specifically program flyers and a module handbook, integrated using Langchain. We benchmark both systems using 90 academic queries categorized by the question difficulty and assess their performance through automatic metrics and blind expert ratings. Our results demonstrate that RAG significantly outperforms the GPT-only approach, particularly for complex questions concerning curriculum and program structure. This research offers valuable insights for higher education institutions seeking to implement reliable and effective AI-powered solutions for student support and information provision.
No overall preference is indicated between human and machine-generated summaries; however, experts with greater familiarity with the Bulletin and higher educational attainment showed a marked preference for the official summaries.
Claudia Biancotti, C. Camassa, Marco Fruzzetti et al.· 0 citations
Before using GenAI models as EdTech tools, their pedagogical suitability should be corroborated. In this paper, we present ShAnEL-2 , a novel multilingual dataset comprising 1,185 student responses to short-answer language learning exercises corrected by teachers. We use ShAnEL-2 to establish an initial benchmark of (1...
Jasper Degraeuwe, Thomas Moerman· International Conference on...· 0 citations
Large Language Models (LLMs) have shown remarkable proficiency on general-purpose tasks, yet their performance often degrades in highly-specialized technical domains. Moreover, little is known about how parametric knowledge of domain-specific terms is encoded within these models. We address this gap by contributing two...
A systematic analysis of existing evaluation frameworks and metrics for RAG-based Question Answering (QA) systems reveals urgent needs for hybrid evaluation frameworks, reasoning path traceability, combined pipeline assessment, and harmlessness evaluation, laying groundwork for the next generation of evaluation methodo...
Nakul Mehta, A. Ojo, Edward Curry· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.