Skip to content
Open access

Maximum Value Attribute-Based Extra Trees, XGBoost, and LightGBM for Early Disease Prediction

Aug 2026 · International Journal of Innovative Science & Technology · 0 citations · 27 references

TL;DR

Experimental results show that the hybrid MVA based models outperforms the traditional classifiers in terms of prediction accuracy, computational efficiency and feature optimization.

Abstract

The evolution of intelligent healthcare systems has increased the importance of machine learning techniques in timely prediction of heart disease by using clinical and medical data. In recent years machine Learning techniques have been used widely in disease prediction systems due to their ability to examine complex medical patterns and also for better diagnostic accuracy. Many machine learning classifiers achieve strong predictive performance but still acquire high computational power and longer execution time during their training and testing phases. The ensemble learning algorithms such as Extra Trees, Extreme Gradient boosting and LightGBM have excellent performance in prediction of data due to their high classification accuracy and robustness. The main problem is that these models often suffer from high iteration counts, longer execution time and the inclusion of irrelevant or redundant features which may affect prediction efficiency and the overall performance of a model. To address these limitations, this research paper introduces a Rough Set Theory which is totally based on hybrid maximum value attribute (MVA). Hybrid MVA is a feature selection method that combines cardinality based ranking for categorical attributes and variance based ranking for continuous attributes in order to choose the most relevant features before training the models and predict heart disease in an efficient manner.The proposed models evaluate three techniques that are hybrid maximum value attribute Extra trees (MVA ET), hybrid maximum value Attribute XGBoost (MVA XGBoost) and hybrid maximum value attribute LightGBM (MVA LightGBM). This hybrid MVA method improves the performance of the model and improves computational efficiency by selecting the most relevant categorical and continuous features and removes irrelevant ones. The models are implemented using Python in Jupyter Notebook. The dataset is obtained from Kaggle platform which is titled as “Synthetic Heart Disease Prediction Dataset” that contains 50000 records and 20 features. The standard performance evaluation metrics including accuracy, precision, recall, F1 score and execution time are used to measure the success of the proposed approach. This report compares standard models against hybrid MVA across various train and test data distributions. The experimental evaluation is performed using accuracy, precision, recall, F1-score, execution time and iteration analysis to measure model effectiveness comprehensively and Experimental results show that the hybrid MVA based models outperforms the traditional classifiers in terms of prediction accuracy, computational efficiency and feature optimization. It reduces the feature dimensionality from 20 to 13 features, i.e., a 35% reduction in the input feature space before model training. For the Extra Trees classifier, the average iteration steps decreased from 46,787 to 31,040, corresponding to a 33.7% reduction in computational complexity. Similarly, training and execution time were also reduced in Hybrid MVA Extra Trees. The Hybrid MVA XGBoost model and Hybrid MVA LightGBM reduced the average iteration steps from 4,560 to 2,964, corresponding to a 35% reduction, while also reducing training and execution times. The Hybrid MVA Extra Trees model achieved an average accuracy of 99.07%, while Hybrid MVA XGBoost and Hybrid MVA LightGBM maintained average accuracies of 99.73% and 99.57%, respectively. Among the proposed approaches, Hybrid MVA XGBoost provided the best balance between predictive performance and computational efficiency.

Read PDF

Similar papers

Jul 2026

An intelligent disease prediction framework integrating machine learning, optimization, and LLM with RAG

A framework that integrates multiple machine learning models and optimization algorithms to enable the early prediction of cardiovascular disease, Parkinson’s disease (PD), and nonalcoholic fatty liver disease (NAFLD) and has high potential for adoption in health-care applications driven by artificial intelligence.

Ping-Huan Kuo, Ping-Chun Tsai, Bang-Yu Chen et al. · 0 citations
Conference Aug 2026

Explainable Machine Learning Framework for Early Detection of Heart Disease Using Feature Selection and Data Balancing

Though technology is extensively applied to medical field, it remains one of the most urgent problems of healthcare sphere as heart diseases are complicated by the interplay of clinical, behavioral, and physiological causes. This research is aimed at suggesting a machine learning framework that can provide accuracy, tr...

Vishal Bharadwaj Meruga, Venkata Reddy Medikonda, Rama Krishna Eluri et al. · 0 citations
Conference Aug 2026

Gradient Boosting Techniques in a Risk-Aware and Explainable Machine Learning Framework for Heart Disease Prediction

The complicated connection between medical risk factors and the serious ramification of misdiagnosis highlights the essential challenge of detecting cardiovascular disease in its early stages. Although most examinations to date have focused on accuracy-centric evaluation, which may not absolutely account for clinical s...

H. Suresh, P. R. · 0 citations
Open access Aug 2026

A Hybrid GA-KNN Framework For Cardiovascular Disease Prediction Using Optimized Clinical Feature Selection

An optimized hybrid approach of the genetic algorithm and K-nearest neighbor method for cardiovascular disease prediction is proposed and it is demonstrated that optimized GA-KNN can achieve both feature dimensions for the initial screening of cardiovascular diseases.

Banibrata Paul, Bhaskar Karn · 0 citations
Conference Aug 2026

An Explainable Machine Learning Framework for Early Detection of Chronic Kidney Disease Using Optimized SVM

The early detection of Chronic Kidney Disease (CKD) is critical in order to commence the right treatment and minimize the chances of the disease from progressing. This paper outlines an interpretable machine learning algorithm for chronic kidney disease classification based on a fine-tuned Support Vector Machine (SVM)....

Parthak Mehra, A. Bist, Devendra Singh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.