Skip to content
Book Open access

NiWo: An Augmentation Framework to Enhance ML Performance and Interpretability for Tabular Data with Class Imbalance

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 10-20 · 0 citations · 15 references

Abstract

Various data augmentation methods have been proposed to address class imbalance in Machine Learning (ML) and Artificial Intelligence tasks across multiple data modalities. For tabular data, augmentation methods must be interpretable so that human decision-makers can audit the process (e.g., which neighborhoods are being augmented and why). This is particularly essential for data from high-stakes domains. Previous studies have integrated explanation tools to derive instance-level weights that drive the augmentation process while maintaining interpretability. However, such an integration is computationally expensive for real-world datasets with thousands of instances due to the complexity of explanation tools. This paper proposes a novel augmentation framework, NiWo, that eliminates the need for explanation tools. NiWo optimizes the weights of influential neighborhood instances within an augmentation budget, i.e., the total number of records to be generated, thus preserving computational efficiency and offering interpretability. It implements two optimization strategies and an ensemble method that activates the most appropriate strategy for a given dataset. NiWo decouples budget allocation from instance generation. The latter can be flexibly replaced to maximize improvement in model performance. This also enables augmentation to be audited and adjusted based on domain-specific requirements within a human-in-the-loop framework. Results of 62 tabular datasets and 7 models show that NiWo outperforms other augmentation methods at enhancing ML performance, especially over datasets with class imbalance and scarce instances.

Read PDF

Similar papers

#large language models Open access Sep 2026

Leakage-Aware LLM Augmentation for Attrition Prediction: A DecisionCentric Evaluation

Employee attrition is a high-cost, asymmetric decision problem plagued by data imbalance. To address this, we propose a leakage-aware, dual-network augmentation framework that integrates the semantic reasoning of Large Language Models (LLMs) with the distributional rigorousness of GAN discriminators. Specifically, we e...

Wei-Quan Liao, Jia-You Xu, Ekaterina A. Panova · 0 citations
Open access 2026

MSTabVAE: Multi-Step Latent Conditional Variational Autoencoder for Imbalanced Tabular Data Synthesis

Recent advances in artificial intelligence have expanded its applications in the financial domain, particularly in fraud detection, a critical task for preventing losses for both customers and institutions. However, fraud detection is challenging due to severe class imbalance, which significantly degrades detection per...

Min-Ji Kang, Hyeryung Jang · 0 citations
Open access Sep 2026

Evaluating Data-Centric Optimization Strategies for Improving Machine Learning Generalization

Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite incre...

Nia Oktaviani, E. Noche, S. Patil et al. · 0 citations
Preprint Aug 2026

Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification

This work explores CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset and demonstrates the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.

Haadia Amjad, Ronald Tetzlaff · 0 citations
Oct 2026

Adaptive Progressive Optimization Ensemble Approach for High-Dimensional Imbalanced Data Classification

High-dimensional imbalanced data presents the problems of massive invalid features and class imbalance, making it arduous for classifiers to gain respectable outcomes. Compared to the individual classifier, classifier ensemble has great potential to elevate the performance. In this paper, an adaptive progressive optimi...

Yuyang Deng, Yu-Hong Xu, Pei-Jie Huang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data

Experiments show that aeSFT identifies useful synthetic data using substantially fewer real test samples than mean-based sequential testing, matches the power of fixed-sample sign-flip testing and the paired $t$-test, while keeping the false-positive rate below the target level.

Zhen-Yu Tao, Wei Xu, Xiao-Hu You et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.