Effective Chunking for Retrieval-Augmented Generation over Structured Institutional Documents
Abstract
Retrieval-Augmented Generation (RAG) has improved domain-specific question answering tasks by grounding responses in an external knowledge base. However, institutional documents differ from general corpora due to the presence of structured content such as tables, flowcharts, fee breakdowns and program codes. Existing RAG pipelines typically employ generic chunking strategies designed for unstructured text, which can fragment structured information and degrade answer quality. This study evaluates four chunking strategies across six configurations. The evaluation uses 65 publicly available structured documents from Universiti Kebangsaan Malaysia (UKM) with 28 Bahasa Melayu queries. Retrieval performance, answer quality and structural integrity are compared across configurations. Retrieval performance is generally similar across configurations, whereas answer quality varies more substantially. Structure-aware chunking achieves the highest ROUGE-L, BLEU and BERTScore F1 values, while semantic chunking achieves the highest table-integrity score. The findings demonstrate that chunking strategies influence RAG objectives in distinct ways, with structure-aware chunking improving answer generation while semantic chunking more effectively preserving table integrity. The results provide an empirical basis for selecting chunking configurations for structured institutional RAG systems.