Santham : A Curated Sanskrit–Tamil Dataset with Anvaya and Segmentation for Building and Evaluating Machine Translation
This work introduces Santham, a curated Sanskrit-Tamil parallel dataset comprising over 90,000 pairs drawn from classical texts such as the Mahābhārata, Rāmāyaṇa, and Bhagavad Gīta, and utilizes anvaya (prose-order reordering) to mitigate the structural complexity of poetic verses.
P. Venkatesh, Ketaki T. S., M. Shetye et al.
· 0 citations