This study presents a systematic empirical investigation of task-adaptive continual pre-training (TAPT), introduced by Gururangan et al., for Turkish language understanding, with a particular focus on the effect of the masked-language-modeling rate.
Abstract
Transformer-based language models have become the standard in Natural Language Processing (NLP). They have surpassed human performance on specific classification tasks such as named-entity recognition, question-answer, text categorization, or generative tasks such as machine translation and summarization. However, since language models are trained with significant general-purpose texts, they may have limitations in their domain-specific knowledge. Techniques such as domain adaptation can be used to improve the models to address this issue. This study presents a systematic empirical investigation of task-adaptive continual pre-training (TAPT), introduced by Gururangan et al., for Turkish language understanding, with a particular focus on the effect of the masked-language-modeling rate. Adaptation is performed in the task-adaptive setting (TAPT), i.e., continual pre-training on the unlabeled text of the target task corpus, without requiring an external domain corpus. We achieved successful results with an average increase of 2.7%. We also addressed various issues and findings related to adaptation.
It is argued that the ability of long context should not only come from increasing the context window, but also from the ability of the model to locate, integrate and reason about important information in long text.
Jun-Hao Wu· Applied and Computational En...· 0 citations
This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies in Bahasa Indonesia and its low...
A. D. Alexander, S. Setiawati· Dinasti Information and Tech...· 0 citations
A corpus of 10,000 Bangla sentences, manually annotated into four functional categories, namely declarative, interrogative, imperative, and exclamatory is introduced, and Experimental results show that TF-IDF consistently outperforms Word2Vec, likely due to its ability to emphasize discriminative lexical cues associate...
Swapnil Kundu Argha, Abdullah Al Shafi, Rowzatul Zannat et al.· 0 citations
A comprehensive review of the evolution of NLP from traditional rule-based approaches to modern transformer models including BERT and GPT demonstrates that NLP continues to transform intelligent systems and is expected to play an increasingly significant role in the development of next-generation AI technologies.
P. Kalaiselvi· International Journal of Eme...· 0 citations
Pretrained language models (PLMs) have established state-of-the-art performance across diverse natural language understanding (NLU) tasks. This study reveals that seman-tic-rich explanations of lexical units can effectively guide PLM learning processes. We propose a novel language understanding enhancement method with...
Tianyi Chen, Yashen Wang, Huan Chang et al.· IEEE/CAA Journal of Automati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.