Skip to content

MaitH 1.0: A Parallel Corpus and Baseline for Low-Resource Maithili-Hindi Translation

2026 · International Conference on Language Resources and Evaluation · pp. 8567-8576 · 0 citations · 29 references
Computer Science

TL;DR

A corpus containing both manually curated and synthetically generated sentences for low-resource Indian languages, such as Maithili is contributed and it is demonstrated that, even with a smaller corpus size, high-quality, task-specific data significantly enhance translation accuracy for low-resource Indian languages, such as Maithili.

View source

Similar papers

Review Open access Sep 2026

Stemming techniques for resource-poor languages: a review of methods, challenges, and applications

A systematic review of stemming techniques across Afro-Asiatic, Indo-Aryan, Turkic, and Uralic language families, with explicit acknowledgment of coverage limitation, points out the strengths and weaknesses of current approaches and provides insights into potential areas for future investigation and developments in lan...

Sanjiban Sekhar Roy, Mohd Anas, Saravanakumar Kandasamy · 0 citations
#natural language process... Preprint Aug 2026

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

Uri Katz, Omer Goldman, Tomasz Limisiewicz et al. · 0 citations
Open access Aug 2026

Dataset Curation for Kalabari NMT System

A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.

O. T. Olise · 0 citations
Preprint Aug 2026

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.

Sujith Pulikodan, A. Basu, J. PavanKumar et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.