Skip to content
Conference

LMSpell: Spell Correction with Pre-Trained Language Models

Aug 2026 · Moratuwa Engineering Research Conference · pp. 503-508 · 0 citations · 39 references

Abstract

Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light on the plight of spell correction for LRLs.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.