Skip to content

Author

R. Mundotiya

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Jul 2026

NERBench-Chhattisgarh: A Multi-Family NER Dataset for Low-Resource Indic Languages

We present NERBench-Chhattisgarh, a gold-standard Named Entity Recognition (NER) dataset covering seven under-resourced languages spoken in Central India: Baigani, Chhattisgarhi, Surgujia, Sadri, Kudukh, Halbi, and Gondi. Addressing the digital divide for tribal languages, our corpus comprises 166,444 annotated tokens across 8,391 sentences, spanning both Indo-Aryan and Dravidian language families. The dataset features high lexical sparsity and a "nature-centric" ontology of 22 entity types based on the CLIA Phase-II schema, capturing culturally specific entities often absent in standard benchmarks. To establish a benchmark for language variety-aware information access, we evaluate three multilingual encoders: mBERT, XLM-RoBERTa, and IndicBERT, using two adaptation strategies: direct parameter-efficient fine-tuning (LoRA) and a Chhattisgarhi-Pivot Adaptive Pre-training (CPAP) approach. Our results show a morphological barrier: while pivot-based adaptation facilitates transfer for distant languages through script alignment, it induces negative transfer in morphologically complex agglutinative languages like Gondi. We publicly release the dataset, code, and adapted model checkpoints. https://github.com/Rajesh-NLP/NER-Chhattisgarh to support future research in inclusive Information Retrieval.

R. Mundotiya · 0 citations