Skip to content
Open access

A Three-Stage Cross-Lingual Knowledge Transfer Approach Based on the XLM-RoBERTa Model for Detecting Fake News in Ukrainian

Aug 2026 · Big Data and Cognitive Computing · Vol 10, pp. 264 · 0 citations · 41 references

TL;DR

This article proposes a three-stage algorithm for training a model to recognize fake news in the Ukrainian language, using the multilingual transformer model XLM-RoBERTa, which solves this problem by utilizing cross-lingual knowledge transfer from English to Ukrainian.

Abstract

In recent years, there has been an increase in the amount of fake news in the media, which is why fact-checking systems are gaining popularity, particularly those that use natural language processing (NLP) to quickly identify and flag fake news. One of the main limitations in the development of such systems is the limited number of datasets containing verified information, which are necessary for the effective training of models. The situation is particularly critical for non-English datasets, specifically those in the Ukrainian language. This article proposes a three-stage algorithm for training a model to recognize fake news in the Ukrainian language. At the core of the proposed approach lies the multilingual transformer model XLM-RoBERTa, which solves this problem by utilizing cross-lingual knowledge transfer from English to Ukrainian. This approach means there is no need to search for a large, high-quality dataset in Ukrainian; instead, a significantly smaller dataset in Ukrainian can be used for the final calibration of the model. The model developed as a result of the experiment proved effective in extreme low-resource scenarios, achieving 90.7% accuracy on just 500 training records and outperforming the baseline model by 9.7%.

Read PDF

Similar papers

Aug 2026

Cross-lingual transfer learning for fake news detection: leveraging high-resource language for low-resource adaptation

A transfer-cum-ensemble learning framework that integrates a task-specific pretrained model (XLM-RoBERTa) with a lightweight Indian-language model (IndicBERT) with a weighted-average attention mechanism is used to combine these sophisticated language models for further enhancing performance.

Garima Thakur, Jyoti Srivastava, Aryan Tyagi et al. · 0 citations
Open access 2026

Integrating Transformer-based and Embedding Models into Rasa NLU for Vietnamese University Support System

This paper presents a Vietnamese university support chatbot developed using the Rasa Natural Language Understanding (NLU) framework, integrating Transformer-based and embedding models, including PhoBERT, FastText, Multilingual BERT (mBERT), and additional baseline methods such as Support Vector Machine (SVM) and Naive Bayes. The system is trained on a domain-specific dataset consisting of 99 intents and 1773 annotated examples covering academic and administrative queries. To ensure reliable evaluation, all models are assessed using 5-fold cross-validation. Experimental results show that PhoBERT achieves the best performance with an average accuracy of approximately 90.5% and an F1-Score of 90.1%, significantly outperforming both traditional machine learning methods and multilingual Transformer models. Among baseline approaches, SVM demonstrates strong performance, highlighting the effectiveness of classical models under limited data conditions. Further analysis using confusion patterns reveals that most misclassifications occur between semantically similar intents, emphasizing the challenges of fine-grained intent classification in Vietnamese. The results confirm that language-specific pretraining plays a crucial role in improving performance in low-resource settings. This study provides an empirical evaluation of multiple modeling approaches under consistent experimental conditions and demonstrates the potential of Transformer-based models for Vietnamese university support systems, while highlighting limitations related to dataset size and intent overlap.

Le Ba Cuong, Le Anh Tien, Huong Van Pham · 0 citations
Open access Aug 2026

Multilingual Fake News Detection Using Machine Learning with Contextual-Based Feature Extraction

The proposed approach provides a simple and efficient solution for multilingual fake news detection in data-scarce environments with ensemble-based classifiers such as Random Forest and Gradient Boosting achieving reliable performance across both languages.

Nikita Garg, Pritam Singh Negi · 0 citations
Open access Jul 2026

Implementation of Transfer Learning for Automatic Summarization in Research Article Synthesis

The findings indicate that BERT-based extractive summarization can support preliminary literature screening, but further improvement is needed through stronger baseline comparison, human evaluation, and redundancy-aware optimization.

Made Hanindia Prami Swari, Puji Lestari Tarigan, Gusti Eka Yuliastuti et al. · 0 citations
Open access Aug 2026

A Distilbert Case-Based Deep Learning Model for Part-of-Speech Tagging for Under-Resourced Kenyan Language: A Case of Dholuo

The results of the experiments reveal that the proposed DistilBERT-based model outperforms the baseline Conditional Random Field and Bidirectional Long Short-Term Memory models with an accuracy 79.05%, precision 80.63%, recall 79.05% and F1-score 79.24%.

Maureen Otieno, L. Wanzare, Calvins Otieno · 0 citations
Open access Sep 2026

Design and Implementation of a Scalable AI-Based Semantic Evaluation System for Hindi Text Using Transformer Models

Evaluating linguistically diverse descriptive answers in a consistent and accurate manner in modern digital education systems is a growing challenge, especially in low-resource languages like Hindi. Traditional lexical and rule-based grading systems cannot adequately reflect the meaning behind the words, negation, paraphrasing, and so on, which leads to low grading reliability. To overcome these limitations, this study proposes an automated evaluation framework with intelligent rule-based linguistic preprocessing and transformer-based deep learning. The framework uses a fine-tuned multilingual BERT (mBERT) model bhavikardeshna/multilingual-bert-base-cased-hindi to provide contextual embeddings, and cosine similarity-based semantic alignment with the model l3cube-pune/hindi-sentence-similarity-sbert is used to provide automated scores. With optimal setting of learning rate = 5×10⁻⁴, batch size = 24 and epochs = 40, accuracy, precision, recall and F1 score of 78.9%, 80.6%, 77.4% and 79.0% respectively is achieved on HindiRC-Data-master dataset (24 passages, 127 question-answer pairs, grades 2-5) which is more than 14% higher than lexical similarity baselines and is better than previous Hindi QA architectures without domain-specific preprocessing pipelines. The suggested system will save about 40% manual grading, and will enable scalable, consistent and repeatable assessment.

Nirja D. Shah, Jyoti Pareek · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.