ITEBIS PGRI Dewantara Academic Service Chatbot Based on IndoBERT and Retrieval-Augmented Generation (RAG)
Abstract
Abstract. Academic administrative services must stay responsive beyond regular hours, yet repetitive student inquiries add to staff workload. Retrieval-Augmented Generation (RAG) with a fine-tuned language model offers a way to automate accurate, round the clock responses in Indonesian. Purpose: This study develops a prototype academic-service chatbot for ITEBIS PGRI Dewantara Jombang, integrating a fine-tuned IndoBERT model with RAG architecture to address limited-service hours and repetitive handling of routine student inquiries. Methods/Study design/approach: The system used 2,517 SQuAD-format question-answer pairs, split into 72% training (1,812), 8% validation (201), and 20% test (504). IndoBERT (indobenchmark/indobert-base-p1) was fine-tuned as an extractive reader for five epochs (batch size 8, learning rate 3 × 10⁻⁵) with early stopping, then converted to ONNX format and INT8-quantized for CPU-based local deployment. The prototype separates a Laravel web interface from a FastAPI backend running the RAG pipeline, retrieving passages from a MySQL knowledge base via FAISS search (top-k = 5, min score 0.30) and catching frequent query-answer pairs in Redis to cut redundant database access and inference latency. Evaluation used black-box testing across five scenarios. Result/Findings: On the 504-pair test set, the model achieved an Exact Match of 71.32% and an F1-Score of 86.63%. All five black-box scenarios passed, including the system’s ability to give database-grounded responses within the tested scenarios rather than fabricated answers, though retrieval for categories with fewer passages (e.g., finance) was more sensitive to question phrasing Novelty/Originality/Value: The novelty lies not in any single element but in combining four components: a fine-tuned IndoBERT reader for Indonesian, an academic-domain RAG architecture, a fully functional end-to-end web prototype rather than model evaluation alone, and an ONNX INT8-quantized reader paired with a Redis caching layer that jointly reduce inference latency and database load a combination not jointly addressed in prior IndoBERT or RAG-based chatbot studies.