Aug 2026· Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi)· 0 citations
TL;DR
It is demonstrated that CVI-validated ensemble GenAI can construct consistent labels for low-resource administrative texts and that IndoBERT provides the strongest and most stable generalization for cooperative supervision classification.
Abstract
Cooperative supervision reports contain complex narrative structures and overlapping administrative terminology, complicating automatic classification into governance, risk profile, financial performance, and capital adequacy. Reliable automation is particularly important for accelerating the analysis of supervisory findings while addressing limited labeled data and imbalanced categories. This study aimed to develop and externally evaluate a text-classification framework combining quantitatively validated Generative Artificial Intelligence (GenAI) labeling with conventional and Transformer-based models. Data comprised 294 preprocessed sentences collected from the Department of Cooperatives, Small and Medium Enterprises, Industry, and Trade of Semarang Regency during 2023–2025. Few-shot annotations were generated using ChatGPT, Gemini, Perplexity, and DeepSeek, and three-model combinations were evaluated using the Content Validity Index (CVI); majority voting from the best combination established ground truth. TF-IDF with Logistic Regression and Support Vector Machine served as baselines, whereas IndoBERT and IndoRoBERTa represented contextual models. Performance was assessed through stratified five-fold cross-validation and external testing on 58 unseen sentences. ChatGPT–Gemini–Perplexity achieved the highest Scale-Level CVI of 0.898. IndoBERT obtained the best cross-validated F1-score of 0.9099, exceeding IndoRoBERTa (0.8217), Logistic Regression (0.8004), and SVM (0.7863). On unseen data, IndoBERT retained an F1-score of 0.862, compared with 0.759 for IndoRoBERTa. These findings demonstrate that CVI-validated ensemble GenAI can construct consistent labels for low-resource administrative texts and that IndoBERT provides the strongest and most stable generalization for cooperative supervision classification. The framework offers a practical basis for scalable annotation and reliable automated support for evidence-based supervisory decision-making.
Treating supervision format as a first-class hyperparameter for multi-task reasoning SFT in large language models—at least in this benchmark-and-model setting—rather than a mere rendering detail is supported.
Nhat Thanh Vu, M. Rashid, Fariza Sabrina· Electronics· 0 citations
This paper presents a Vietnamese university support chatbot developed using the Rasa Natural Language Understanding (NLU) framework, integrating Transformer-based and embedding models, including PhoBERT, FastText, Multilingual BERT (mBERT), and additional baseline methods such as Support Vector Machine (SVM) and Naive Bayes. The system is trained on a domain-specific dataset consisting of 99 intents and 1773 annotated examples covering academic and administrative queries. To ensure reliable evaluation, all models are assessed using 5-fold cross-validation. Experimental results show that PhoBERT achieves the best performance with an average accuracy of approximately 90.5% and an F1-Score of 90.1%, significantly outperforming both traditional machine learning methods and multilingual Transformer models. Among baseline approaches, SVM demonstrates strong performance, highlighting the effectiveness of classical models under limited data conditions. Further analysis using confusion patterns reveals that most misclassifications occur between semantically similar intents, emphasizing the challenges of fine-grained intent classification in Vietnamese. The results confirm that language-specific pretraining plays a crucial role in improving performance in low-resource settings. This study provides an empirical evaluation of multiple modeling approaches under consistent experimental conditions and demonstrates the potential of Transformer-based models for Vietnamese university support systems, while highlighting limitations related to dataset size and intent overlap.
Le Ba Cuong, Le Anh Tien, Huong Van Pham· International Journal of Inf...· 0 citations
The results demonstrate that superior predictive performance does not necessarily correspond to higher explanation faithfulness or stronger cross-domain stability, and highlight the importance of jointly evaluating predictive performance, explanation faithfulness, and explanation robustness when developing trustworthy transformer-based NLP systems.
Dony Bahtera Firmawan, B. Darnoto· Journal of Computing Theorie...· 1 citation
Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%.
Batyr Sharimbayev, S. Kadyrov· Journal of Advances in Infor...· 0 citations
It is taken as initial evidence for market time series as an input modality in financial text classification on the task of classifying sentences from Federal Reserve communication as hawkish, dovish, or neutral.
Michael Schlee, Fabian Lukassen, Christoph Weisser· 0 citations
This study evaluated GPT-4o-based data augmentation for imbalanced multiclass sentiment classification of GoPay user reviews using IndoBERT-LoRA. The main problem addressed in this study was the limited representation of minority sentiment classes, particularly the neutral class, which could reduce the model’s ability to recognize all sentiment categories proportionally. The dataset consisted of 16,955 Google Play Store reviews that were manually labeled into positive, neutral, and negative classes. Two augmentation strategies were compared, namely prompt-based augmentation and fine-tuning augmentation. The generated synthetic data were evaluated using novelty, diversity, duplication, and manual validation of sampled reviews before being incorporated into the training data. The IndoBERT-LoRA model was trained under four scenarios: baseline, class weighting, prompt-based augmentation, and fine-tuning augmentation. The results showed that fine-tuning produced better lexical-level quality indicators, as indicated by more stable novelty and diversity scores and a lower duplication rate. Both augmentation strategies improved macro recall and macro F1-score compared with the baseline and class weighting scenario. The largest improvement occurred in the neutral class, where recall increased from 0.5633 to 0.7801 with prompt-based augmentation and to 0.8133 with fine-tuning augmentation. These findings indicate that GPT-4o-based augmentation improved minority-class recognition, although the improvement involved a trade-off between recall, precision, and implementation cost.
M. Yusran, F. Afendi, Anwar Fitrianto· International Journal of Adv...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.