Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· pp. 10-20· 0 citations· 15 references
Abstract
Various data augmentation methods have been proposed to address class imbalance in Machine Learning (ML) and Artificial Intelligence tasks across multiple data modalities. For tabular data, augmentation methods must be interpretable so that human decision-makers can audit the process (e.g., which neighborhoods are being augmented and why). This is particularly essential for data from high-stakes domains. Previous studies have integrated explanation tools to derive instance-level weights that drive the augmentation process while maintaining interpretability. However, such an integration is computationally expensive for real-world datasets with thousands of instances due to the complexity of explanation tools. This paper proposes a novel augmentation framework, NiWo, that eliminates the need for explanation tools. NiWo optimizes the weights of influential neighborhood instances within an augmentation budget, i.e., the total number of records to be generated, thus preserving computational efficiency and offering interpretability. It implements two optimization strategies and an ensemble method that activates the most appropriate strategy for a given dataset. NiWo decouples budget allocation from instance generation. The latter can be flexibly replaced to maximize improvement in model performance. This also enables augmentation to be audited and adjusted based on domain-specific requirements within a human-in-the-loop framework. Results of 62 tabular datasets and 7 models show that NiWo outperforms other augmentation methods at enhancing ML performance, especially over datasets with class imbalance and scarce instances.
Employee attrition is a high-cost, asymmetric decision problem plagued by data imbalance. To address this, we propose a leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models (LLMs) with the distributional rigorousness of GAN discriminators. Specifically, we e...
Recent advances in artificial intelligence have expanded its applications in the financial domain, particularly in fraud detection, a critical task for preventing losses for both customers and institutions. However, fraud detection is challenging due to severe class imbalance, which significantly degrades detection per...
Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite incre...
Nia Oktaviani, E. Noche, S. Patil et al.· Journal of Data Science· 0 citations
This work explores CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset and demonstrates the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.
High-dimensional imbalanced data presents the problems of massive invalid features and class imbalance, making it arduous for classifiers to gain respectable outcomes. Compared to the individual classifier, classifier ensemble has great potential to elevate the performance. In this paper, an adaptive progressive optimi...
Yuyang Deng, Yu-Hong Xu, Pei-Jie Huang et al.· IEEE Transactions on Knowled...· 0 citations
Experiments show that aeSFT identifies useful synthetic data using substantially fewer real test samples than mean-based sequential testing, matches the power of fixed-sample sign-flip testing and the paired $t$-test, while keeping the false-positive rate below the target level.
Zhen-Yu Tao, Wei Xu, Xiao-Hu You et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.