Skip to content

Santham : A Curated Sanskrit–Tamil Dataset with Anvaya and Segmentation for Building and Evaluating Machine Translation

· 0 citations · 20 references

TL;DR

This work introduces Santham, a curated Sanskrit-Tamil parallel dataset comprising over 90,000 pairs drawn from classical texts such as the Mahābhārata, Rāmāyaṇa, and Bhagavad Gīta, and utilizes anvaya (prose-order reordering) to mitigate the structural complexity of poetic verses.

View source

Similar papers

Open access Aug 2026

Dataset Curation for Kalabari NMT System

A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.

O. T. Olise · 0 citations
Jul 2026

Systematic Analysis of Large Language Models and Transformer-Based Machine Translation for English-Tamil and Tamil-English Across Diverse Datasets

The findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.

S. Sriharshaa, Sangeetha Sivanesan, S. Nirmala · 0 citations
Preprint Aug 2026

Padamitra: Grounded Glossary Generation for Classical Sanskrit

We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,3...

Manoj Balaji Jagadeeshan, Sai Pragnaan Marala, Pawan Goyal · 0 citations
Open access Aug 2026

Low-Resource Machine Translation of Yoruba Text to Nigerian Pidgin Using NLLB-200 and Parameter-Efficient Fine-Tuning

Machine translation for low-resource and structurally divergent languages remains a significant challenge, particularly when mapping highly tonal languages to contact languages with fluid orthographies. This study presents the development of a translation system for Yoruba to Nigerian Pidgin, which addresses a critical...

Oyebola Akande, G. Opateye · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.