Skip to content

IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages

Jul 2026 · arXiv.org · Vol abs/2607.23242 · 0 citations · 20 references
Computer Science

TL;DR

IndicTalk is presented, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages and will be released to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages.

Abstract

Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and human evaluations demonstrate that IndicTalk produces fluent, coherent, and naturally code-mixed conversations across both script variants. We will release IndicTalk to support the development and evaluation of multilingual conversational AI for underrepresented Indic languages. The dataset is available at: https://huggingface.co/datasets/LingoIITGN/IndicTalk .

View source

Similar papers

Preprint Aug 2026

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.

Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al. · 0 citations

TEIDAN: A Multilingual Multiparty Dialogue Corpus

This paper describes the collection design, participants, recording setup, transcription format, and corpus statistics, and provides preliminary analyses to illustrate how TEIDAN can support research on turn-taking, addressee recognition, and multimodal grounding in human-human and human-agent interaction.

Taiga Mori, K. Inoue, Mikey Elmers et al. · 1 citation
#natural language process... Preprint Sep 2026

TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired Audio

TRILOGUE (TRIlingual spoken diaLOGUE fact-checking) is introduced, a large-scale trilingual benchmark of source-grounded spoken dialogues in English, Russian, and Kazakh that supports claim check-worthiness detection, source-article evidence retrieval, and claim verification with claim-only, gold-evidence, and retrieve...

Chaewan Chun, Meruyert Aristombayeva, Jiyoung Choi et al. · 0 citations

LLM performance on multi-interlocutor NLI tasks

It is concluded that genuine multi-speaker dialogue inference remains an unsolved problem for current LLMs under standard prompting strategies.

Unknown authors · 0 citations
#artificial intelligence Preprint Sep 2026

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu...

Farah Adeeba, A. Khan, Rajesh Bhatt et al. · 0 citations
#natural language process... Preprint Sep 2026

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

This work proposes two approaches of code-mixed generation using parallel sentences of three languages and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages for language identification in code-mixed settings.

Pruthwik Mishra, Rudra H. Trivedi, Avi Patel et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.