Skip to content
Conference Open access

Hybrid Retrieval for Indic Language Question Answering using an Improved RAG Pipeline

Sep 2026 · Engineering & Technology · 0 citations

Abstract

Although Large Language Models (LLMs) have progressed significantly, Retrieval-Augmented Generation (RAG) still falls short for under-resourced Indic languages. Existing embedding systems, built primarily for English, face issues with script mismatches and limited word coverage, leading to near-complete retrieval failures in mixed Hinglish contexts. This paper introduces Hybrid Indic-RAG, a simple framework to reduce meaning shifts. It uses a targeted word-matching layer to connect Romanised Hinglish with standard Devanagari Hindi, paired with a two-part scoring system that favours exact word matches over fuzzy semantic signals. Tests in fields like healthcare and books show it cuts retrieval errors by 45% and lifts fact accuracy to 75%, well ahead of standard RAG methods. Our approach offers a low-cost way to build stable, error-free question-answering tools for India's language ecosystem.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.

Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak · 1 citation
Preprint Aug 2026

BengaliMCQ: Automatic Generation and Answer Prediction of Academic Multiple-Choice Questions in a Low-Resource Language

A structure-aware RAG framework is proposed that models Bengali textbooks as hierarchical graphs and uses a contrastively trained graph neural network to retrieve a small set of relevant passages, enabling topic-specific multiple-choice question (MCQ) generation and in-domain answer prediction.

Abu Tarabin Surzo, A. K. M. Nihalul Kabir, Sm Azmain Faysal et al. · 0 citations
#machine learning Preprint Sep 2026

Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints

Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local deployment setting. The primary benchmar...

Imtiaz Ul Hassan, Öykü Akbulut, Onur Kaya et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.