Skip to content

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

Jul 2026 · arXiv.org · Vol abs/2607.23322 · 0 citations · 42 references
Computer Science

TL;DR

Model quality does not increase monotonically with data curation, a result it is reported that data-quality gains do not increase monotonically with data curation.

Abstract

Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.

View source

Similar papers

Conference Jul 2026

An On-Premise Multilingual Academic Chatbot using Retrieval-Augmented Generation and Context-Aware Memory for University Assistance

Universities now use Large Language Models (LLMs) to transform their processes for managing student information. The paper introduces an upgraded chatbot system for Narasaraopeta Engineering College (NEC) which extends previous on-premise LLM chatbot research by providing four new functions. The system uses (1) Retriev...

M. Yaswanth, Kopparapu Sai Amar Durgesh, Mogili Harsha Vardhan et al. · 0 citations
Review Sep 2026

Natural Language Processing and Large Language Models: AI for Education, Communication, and Digital Content Analysis

Natural language processing (NLP) has progressed from narrow, task-specific statistical models to large language models (LLMs) capable of generating fluent, contextually coherent text across virtually unlimited domains, a transition that has reshaped three previously distinct application areas simultaneously: education...

R. Rakesh · 0 citations
Open access Aug 2026

ChatGPT vs. Gemini: Which model generates more comprehensible texts and more solvable tasks?

Large language models (LLMs), which are computer programs designed to understand and generate human language, are increasingly common in foreign language teaching. These models can reduce the time and effort teachers need to put in, and they help students learn a foreign language on their own, without frustration or fe...

Lukas Paun · 1 citation
Open access Jul 2026

Direct Use of Corpora in Japanese Language Education: Implementing DDL at the Upper-Intermediate Level

A four-year project integrated direct corpus consultation through NINJAL corpora accessed via Chūnagon, with the dual aim of supporting language development and fostering research-oriented skills for working with Japanese primary sources, offering practical insights into the feasibility of DDL in Japanese language educ...

Giuseppe Pappalardo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.