Skip to content

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

Jul 2026 · arXiv.org · Vol abs/2607.23319 · 0 citations · 42 references
Computer Science

TL;DR

BHARATI is presented, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages with native script support for all languages, and three successive tokenizer versions are described.

Abstract

Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fertility analysis demonstrates that v3 averages 2.6 tokens per Indian Knowledge System (IKS) technical term, compared to 5.25 tokens per term with GPT-2's tokenizer and 3.75 tokens with the multilingual SentencePiece baseline, with the largest gains on a set of reserved IKS terms that are represented as single tokens by construction. On a held-out test set of 490 IKS-domain sentences (70 per language across seven languages, released with the measurement script), v3 reduces sequence length by roughly 90% relative to GPT-2 and byte-level encoding (which lack native Indic subwords) and by approximately 25% relative to the mBART-50 multilingual baseline, averaged across the six Indic languages, directly translating to increased effective context length for downstream language models. The tokenizer models (32,000 vocabulary), training scripts, and evaluation benchmarks are released under open licenses.

View source

Similar papers

#natural language process... Preprint Sep 2026

A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation

Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance....

V. HariKrishnanK, Sudarsun Santhiappan · 0 citations
Jul 2026

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

This work quantifies the tokenizer tax for Indian languages using the FLORES-200 parallel corpus, measuring tokenization fertility across six widely used tokenizers and fourteen languages, and identifies the primary mechanism behind this disparity: failed byte-pair merges that leave text fragmented into single-byte tok...

Priyanshu Srivastava · 0 citations
Preprint Aug 2026

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

Byte-level BPE tokenizers that use the HuggingFace ByteLevel pre-tokenizer inherit GPT-2's word regex, where a word is defined as \p{L}+, one or more Unicode letters, which splits each word at every vowel sign, and formalises this effect as a training-free lower bound on fertility.

S. Regmi, Siddhartha Pudasaini, Chetan Phakami Pun · 2 citations
Aug 2026

Analysis of the Efficiency of Subword Tokenizers in a Low-Resource Linguistic Environment: Implementation Experience for the Tajik Language

The experimental results revealed the strengths and weaknesses of various approaches to subword segmentation and identified the most effective tokenization strategies under the conditions of the morphological complexity of the Tajik language.

M. Arabov, S. S. Khaybullina · 0 citations
Open access Sep 2026

Jawhar: Optimized Morphological Analysis and Contextual Reranking for Arabic Part-of-Speech Tagging

Part-of-speech (POS) tagging in Arabic is hard because its rich root-and-pattern morphology and the absence of short vowels make one unvoweled string compatible with many categories. This paper presents Jawhar, a hybrid framework that couples a high-performance morphological analyser with contextual reranking using a p...

Mohamed Bouzahir, A. A. Abdelouahad, M. Nabil · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.