Skip to content
Open access

தமிழ் வட்டார மொழிகளை AI மொழி மாதிரிகள் கையாளும் திறன்: கொங்கு வட்டாரத் தமிழை மையமாகக் கொண்ட ஆய்வு

Aug 2026 · Tamilmanam International Research Journal of Tamil Studies · 0 citations

TL;DR

A small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary and preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions.

Abstract

While Large Language Models (LLMs) demonstrate significant capabilities in standard Tamil, the extent to which they accurately understand and translate Kongu regional Tamil—spoken in the Kongu region including Coimbatore, Tiruppur, Erode, Salem, and Karur—remains an insufficiently explored area. This article examines the dialect-handling capabilities of AI language models, focusing on Kongu Tamil. For this purpose, a small-scale pilot Kongu Tamil benchmark dataset was curated by the author based on a verified Kongu dialect vocabulary. Using this dataset, zero-shot and few-shot experiments were conducted with a publicly accessible Large Language Model (Claude), and the responses were evaluated based on the author's linguistic knowledge, with errors categorized accordingly. Additionally, citing results from previously published studies comparing various LLMs (Gemini, ChatGPT, Claude), this paper proposes a comprehensive methodological plan for a complete multi-model comparative study for Kongu Tamil, along with a human expert evaluation protocol. Preliminary findings indicate that while the language model handles lexical dialectal variations reasonably well, significant errors occur in cultural nuances, syntactic variations, and idiomatic expressions. This study underscores the need to build large-scale, human-expert-verified benchmark datasets for Tamil dialects.

Read PDF

Similar papers

Open access Aug 2026

கணினியில் தமிழ் வளர்ச்சியும் நவீன இயல்மொழி செயலாக்கத் தொழில்நுட்பங்களும்: ஓர் விரிவான ஆய்வறிக்கை

With the rapid growth of computer technology, the application of the Tamil language has expanded globally. Overcoming the early limitations of 8-bit ASCII encoding systems, the introduction of 16-bit encoding standards such as Unicode and TACE16 has enabled the seamless use of Tamil characters across all computer and m...

திருமதி சு.கலைமதி, திருமதி ரா அனிதா · 0 citations
Open access Aug 2026

செயற்கை நுண்ணறிவு காலத்தில் தமிழ் மொழி மற்றும் பண்பாட்டு உள்ளடக்கங்களின் டிஜிட்டல் வணிகமயமாக்கல்: தொழில்முனைவு வாய்ப்புகள் மற்றும் சவால்கள்

How Tamil culture is being commercialized across diverse platforms such as Tamil Natural Language Processing (NLP), voice technology, machine translation, digital publishing, e-commerce platforms, social media content creation, and Educational Technology (EdTech) is analyzed.

V. K. Veerakumar, G. Sasikaladevi, D. Rajakumari · 0 citations
#natural language process... Preprint Aug 2026

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish, is presented, providing both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

Uri Katz, Omer Goldman, Tomasz Limisiewicz et al. · 0 citations
#natural language process... Preprint Sep 2026

CantoneseLLM v2: Reasoning in a Low-Resource Language

Evaluation across the training stages shows that chat-vector merging transfers instruction following but preserves the donor model's reasoning language, while SFT with limited Cantonese reasoning data substantially shortens or removes reasoning traces and reduces benchmark performance.

Tsz-Chung Cheng, Chung-Shing Cheng, Chaak-ming Lau et al. · 0 citations
Open access Sep 2026

பெருமாள் முருகன் நாவல்களில் கொங்கு வட்டார மொழியும் உள்ளூர் அறிவு மரபுகளும்: Corpus Linguistics மற்றும் Digital Humanities அடிப்படையிலான ஆய்வு

This article examines the Kongu regional linguistic elements found in these novels and the local knowledge traditions they reveal and the need for and possibilities of documenting regional dialect literary works through digital humanities tools using Corpus Linguistics and Digital Humanities methodologies.

முனைவர் இல. பூவலிங்கம் · 0 citations

BankGPT: the use of large language models in official communications 1

No overall preference is indicated between human and machine-generated summaries; however, experts with greater familiarity with the Bulletin and higher educational attainment showed a marked preference for the official summaries.

Claudia Biancotti, C. Camassa, Marco Fruzzetti et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.