Hypertension remains a significant global risk factor for cardiovascular disease and related mortality, necessitating reliable early risk prediction models. Although boosting algorithms have demonstrated strong performance in structured medical data, limited studies have examined their consistency across heterogeneous datasets. This study aims to evaluate the cross-dataset performance and stability of three boosting models, such as XGBoost, LightGBM, and CatBoost, for hypertension prediction under multiple train–test split ratios. Two independent structured datasets were analyzed using 60:40, 70:30, 80:20, and 90:10 splits. To identify the optimal hyperparameters, grid search was performed using repeated stratified 5-fold cross-validation with three repetitions. Model effectiveness was measured using the evaluation metrics of accuracy, precision, recall, F1-score, and AUC. Results show that Dataset 1 gained consistently high predictive performance (accuracy > 0.98; AUC ≈ 1.00), indicating strong and well-separated predictive signals, whereas Dataset 2 demonstrated substantially lower discriminative ability (accuracy ≈ 0.71–0.72; AUC ≈ 0.50), suggesting limited predictive structure. Across both datasets, CatBoost consistently obtained the highest accuracy, particularly at the 90:10 split ratio. These findings demonstrate that dataset characteristics critically determine model effectiveness and that among the evaluated boosting algorithms, CatBoost delivered the strongest overall predictive performance.
Bety Wulan Sari, D. Murtiningsih, Donni Prabowo et al.· Jurnal RESTI (Rekayasa Siste...· 0 citations
The social media platform X has become an important channel for the public to express opinions and share experiences regarding public services, including Trans Jogja. User-generated content from this platform provides valuable insights into public perceptions of service quality. However, because these data consist of unstructured text, sentiment classification techniques based on machine learning are required to analyze them effectively. This study aims to compare the performance of several machine learning algorithms for sentiment classification, including Naïve Bayes, Support Vector Machine (SVM), Random Forest, Neural Network, Logistic Regression, and Decision Tree, in classifying user sentiment toward Trans Jogja on the X platform. Data were collected through a web crawling process using Tweet Harvest with keywords related to Trans Jogja, covering the period from January 1, 2025, to June 10, 2026, resulting in a dataset of 3,035 tweets. The preprocessing stage included data cleaning, case folding, tokenization, normalization, stopword removal, and stemming. Text representation was performed using the Term Frequency–Inverse Document Frequency (TF-IDF) method. The dataset was then divided into training and testing sets using five train–test split ratios: 90:10, 85:15, 80:20, 75:25, and 70:30. Model performance was evaluated using a confusion matrix and the corresponding accuracy, precision, recall, and F1-score metrics. The experimental results demonstrate that the Support Vector Machine (SVM) consistently outperformed the other algorithms across different data split ratios. At the 85:15 train–test split, the SVM achieved an accuracy of 91%, precision of 91%, recall of 91%, and an F1-score of 91%, indicating that it is the most effective algorithm for sentiment classification of Trans Jogja users on the X platform.
Putri Muryanti Setyowati, Y. Pristyanto, Arif Nur Rohman· SISTEMASI· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.