Jul 2026· GECCO Companion· pp. 1166-1174· 0 citations· 28 references
Computer Science
TL;DR
A clustering-based knowledge base is proposed that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity.
Abstract
Mining classification rules for the minority class is challenging not only because positive examples are rare, but because they concentrate in small, geometrically irregular subregions of feature space that unguided search methods systematically miss. We propose a clustering-based knowledge base that captures positive-class structure before search begins and uses it to focus both initialization and neighborhood exploration in Moca-I, a multi-objective local-search algorithm that mines interpretable rule sets by trading off minority-class recall, precision, and complexity. Rather than sampling attribute conditions blindly from the full discretized space, the knowledge base clusters positive-class instances, extracts per-cluster attribute ranges, and aligns them with the algorithm's discretization—seeding the initial archive with minority-class-informed prototypes and restricting neighborhood operators to locally relevant regions. We instantiate this framework with two clustering methods: Self-Organizing Maps (MOCA-ISOM), which preserve the topological structure of the positive-class manifold, and K-Means (MOCA-IKM), a centroid-based baseline. Evaluated on 19 imbalanced benchmark datasets, MOCA-ISOM achieves statistically significant F-measure improvements on eight datasets and Hypervolume improvements on nine, with gains up to +17% absolute on high-dimensional data. MOCA-IKM shows comparable average performance but exhibits five statistically significant degradations.
This study proposes complexity-guided cluster-based adaptive oversampling (CCAS), a novel resampling technique for low- to mid-sized imbalanced datasets that outperforms the benchmark methods across metric-classifier combinations.
Abhilash Panda, U. Chattaraj· Sādhanā· 0 citations
Fuzzy Random Forest is particularly suited for applications requiring high precision in minority class identification and interpretable fuzzy decision rules, such as medical diagnosis, fraud detection, and credit risk assessment where regulatory compliance demands transparent decision-making.
James Omusula Atsali· Asian Journal of Probability...· 0 citations
Text categorization remains a challenging task due to the inherent ambiguity of natural language, class overlap, and imbalanced topic distributions. Traditional Fuzzy C-Means (FCM) clustering, although widely used for soft text classification, is highly sensitive to initialization and tends to favor dense or majority c...
Michael Loki, Agnes Mindila, W. Mwangi· Scientific Reports· 0 citations
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
Generalized Category Discovery (GCD) requires a model to recognize labeled seen classes while discovering unlabeled novel classes from partially annotated data. This paper argues that a key difficulty comes from bias that accumulates during optimization. Two effects appear repeatedly in practice: feature response imbal...
Feng-Xiang Su, Chao-Fan Dai, Hao-Hao Zhou et al.· 2026 12th International Conf...· 0 citations
Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.
Saifur Rahman Mazumder, Feng Yu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.