This work introduces Santham, a curated Sanskrit-Tamil parallel dataset comprising over 90,000 pairs drawn from classical texts such as the Mahābhārata, Rāmāyaṇa, and Bhagavad Gīta, and utilizes anvaya (prose-order reordering) to mitigate the structural complexity of poetic verses.
A transferable curation framework for endangered languages, the first sizable Kalabari-English parallel corpus, and baseline experiments that reveal both the promise and the hallucination pitfalls of training on highly constrained, domain-specific data are contributed.
O. T. Olise· WORLD JOURNAL OF INNOVATION...· 0 citations
The findings show that the quality of the datasets and their alignment with the domain will greatly affect the performance of the model, that attention-based mechanisms can aid in explain ability, and that few-shot large language models can still produce structurally coherent translations of Tamil.
S. Sriharshaa, Sangeetha Sivanesan, S. Nirmala· arXiv.org· 0 citations
We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,3...
Manoj Balaji Jagadeeshan, Sai Pragnaan Marala, Pawan Goyal· 0 citations
Machine translation for low-resource and structurally divergent languages remains a significant challenge, particularly when mapping highly tonal languages to contact languages with fluid orthographies. This study presents the development of a translation system for Yoruba to Nigerian Pidgin, which addresses a critical...
Oyebola Akande, G. Opateye· Cureus Journal of Computer S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.