Skip to content

RUBRIC: Realism-Utility Balanced Ranking for Imbalanced Classification

Jul 2026 · arXiv.org · Vol abs/2607.09816 · 1 citation · 38 references
Computer Science

TL;DR

This work introduces RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem and shows that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators.

Abstract

Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates that distort decision boundaries or introduce artifacts, leading to overfitting and degraded generalization. In this work, we introduce RUBRIC, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem. RUBRIC ranks candidates using a realism-utility trade-off: realism is quantified by a learned discriminator that distinguishes real samples from synthetic samples, while utility captures proximity to the decision boundary through a concave margin-based scoring function. We show that, under mild regularity conditions, the proposed filtering strategy monotonically tightens the generalization bound for margin-based classifiers by jointly reducing distribution shift and suppressing near-negative tail contributions. Through extensive experiments on credit-card fraud detection and other imbalanced benchmarks, we demonstrate that RUBRIC improves F1-macro and recall while maintaining comparable ROC-AUC across several generators. We also provide explicit lambda-sensitivity analysis to show how users can recover AUPRC when ranking quality is prioritized.

View source

Similar papers

Open access 2026

NADOS: Reliability-Gated Difficulty-Aware Oversampling for Noisy Imbalanced Classification

Noise-Aware Adaptive-Difficulty Oversampling (NADOS) is proposed, which separates the assessment of seed trustworthiness from the allocation of synthesis effort and provides a practical strategy for noisy imbalanced classification.

Ednel Ashraff Misran, Syahid Anuar, Adam Mohd Khairuddin · 0 citations
Open access Sep 2026

Adaptive Weighting–Synthetic Minority Oversampling Technique

Class imbalance is prevalent in real-world datasets. Minority samples are far fewer than majority samples. Traditional classifier design typically assumes balanced data, which causes classifiers to favor the majority class when faced with imbalanced datasets. Thus, there are high misclassification costs for minority cl...

Shen Yan, Hai-Feng Guo, Xiao-Ming Su · 0 citations
Jul 2026

GAN-powered oversampling with strengthened sample distribution for class overlapping imbalanced data

This study introduces a novel approach, named SDCGAN, which oversamples with a Conditional Generative Adversarial Network (CGAN) where the generator is fed with strengthened distribution information extracted from a curated set of minority samples.

Yu-Lin Zhang, Xiao-Zhe Wang, Kai-Wen Xue et al. · 0 citations
Open access Jul 2026

SafeVAE-GAN: A Novel Re-sampling Approach for Multi-Class Imbalanced Data

Multi-class imbalanced datasets are ubiquitous in domains like medical diagnostics, fraud detection, and learning performance classification, where minority classes are critical but underrepresented, and class overlap introduces noise and ambiguous boundaries. Prior research has explored oversampling and undersamplin...

Nhiem Ba Nguyen, Sinh Van Nguyen, B. Nguyễn · 0 citations
#machine learning Preprint Aug 2026

RoBell-RVFL: A Robust Generalized Bell Random Vector Functional Link Network

RoBell-RVFL is proposed, a robust and lightweight generalized bell random vector functional link network that redefines how randomized models handle class imbalance and noisy data and achieves adaptive control over sample contributions without sacrificing the closed-form learning efficiency of RVFL networks.

A. Rahaman, A. Quadir, M. Tanveer · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.