A Computational Approach to Plagiarism Detection in Kannada Texts Using Sentence, Bigram, and Trigram Similarity Models
This paper documents a Python implementation for comparing Kannada Unicode texts using sentence-level sequence similarity and character- and word-level bigram and trigram similarities. The application normalizes text with NFKC (Normalization Form Compatibility Composition), collapses whitespace, segments sentences using Kannada danda and common punctuation, computes Jaccard and cosine n-gram scores, and produces an interpretable report. In the supplied combined-mode run, Kadyanata contained 1,388 segmented source sentences and Karavaliya Saviradondu Daivagalu contained 219 target sentences. At a sentence threshold of 0.60, 138 target sentences matched their best source sentence, giving a sentence-match rate of 63.01%. Character bigram and trigram cosine similarities were 96.05% and 88.89%, while word bigram and trigram cosine similarities were 15.43% and 6.25%. The mean of the eight implemented n-gram scores was 36.08%. The code-defined combined score, 40% sentence-match rate plus 60% mean n-gram similarity, was 46.85%.