The impact of balancing real and SD is examined and some recommendations that researchers can further utilise to improve the ML model’s training process are provided and which approaches to adopt are considered.
Abstract
Data serves as the foundation of contemporary artificial intelligence (AI) systems, yet ethical and practical constraints often limit the availability and usability of real-world datasets. Synthetic data (SD) has emerged as a valuable solution, enabling the development and training of AI models without compromising privacy standards or ethical guidelines. However, many challenges remain to address, from generating high-quality SD using various approaches to investigating the impacts of data training on machine learning (ML) models. This study examines the impact of balancing real and SD and provides some recommendations that researchers can further utilise to improve the ML model’s training process. Three datasets, Mobile Health (MHealth), High-Energy Physics Mass (HEPMass), and US Company Bankruptcy Prediction (UCBP), were pre-processed to ensure compatibility and used as the basis for SD generation using Conditional Tabular Generative Adversarial Networks (CTGAN) and Tabular Variational Autoencoders (TVAE). The study employed a hybrid data generation approach, splitting the training data into varying proportions of real and SD, with performance evaluated through 5-fold cross-validation on Decision Tree (DT), Gaussian Naive Bayes (GNB), and Linear Support Vector (L-SVM) ML models. The results indicate the optimal balance of real and SD for maximising model performance. Analysis of benchmarking results across three datasets shows that combining 30% real data with 70% CTGAN-generated synthetic data achieves the highest accuracy and overall model performance. In contrast, when using TVAE-generated data, a 20% real and 80% synthetic split is recommended to maintain similar performance. Additional experiments conducted on 10% reduced subsets showed that while the primary trends persisted under limited-data conditions, the optimal real-to-synthetic data ratio became more sensitive to the specific dataset. Statistical significance of the model performance differences was further confirmed using paired t-tests across all evaluated mixing ratios. This research analyses these results and considers the state of the art to recommend how further synthetic data can be useful with machine learning models and which approaches to adopt.
US healthcare systems struggle with hospital readmissions, especially within 30 days of release. ML-based prediction of hospital readmission is a significant milestone for healthcare professionals to efficiently manage healthcare quality, premature readmissions, post-discharge follow-ups, incur large costs and identify...
Methun Kamruzzaman, Sujoy Saha, Md Nazmul Alam Bhuiyan et al.· Frontiers in Computer Scienc...· 0 citations
NiWo optimizes the weights of influential neighborhood instances within an augmentation budget, thus preserving computational efficiency and offering interpretability, and outperforms other augmentation methods at enhancing ML performance, especially over datasets with class imbalance and scarce instances.
Asif Ahmed, Sakhawat Hossain Saimon, Jianhua Ruan et al.· Proceedings of the 32nd ACM...· 0 citations
Class imbalance is a critical challenge in the classification of tabular data, since it affects the diagnostic capacity of models in domains such as health and finance. This research compares four synthetic data generation paradigms: traditional interpolation (SMOTE-NC), deep generative models (CTGAN and TVAE), and a h...
Jhonatan Esquivel, Christian Humpiri, Jose M. Vega et al.· International Journal of Adv...· 0 citations
Experiments show that aeSFT identifies useful synthetic data using substantially fewer real test samples than mean-based sequential testing, matches the power of fixed-sample sign-flip testing and the paired $t$-test, while keeping the false-positive rate below the target level.
Zhen-Yu Tao, Wei Xu, Xiao-Hu You et al.· 0 citations
The findings indicate that while modern generative models can produce highly realistic and analytically useful datasets, persistent challenges remain, including the lack of standardized benchmarking protocols, utility–privacy trade-offs, privacy leakage risks, bias amplification, limited explainability, and governance...
N. Emran, Ruhaila Maskat, Abdulrazzak Ali· International journal of res...· 0 citations
Cardiovascular diseases (CVDs) remain the foremost cause of mortality worldwide, claiming approximately 17.9 million lives annually and representing 32% of all global deaths. Timely and accurate detection of heart disease is therefore of paramount importance. This paper presents a narrative survey, synthesising 57 peer...