Skip to content

A Leakage-Free Stacked Ensemble Method for Multiclass Classification

Jul 2026 · arXiv.org · Vol abs/2607.22081 · 0 citations · 30 references
Computer Science

TL;DR

LFS-FRAME is proposed, a Leakage-Free Stacked ensemble framework that integrates functional learning using Kolmogorov-Arnold Networks (KAN) and rule-based learning via XGBoost and rule-based learning via XGBoost for robust multiclass classification.

Abstract

Multiclass classification is a fundamental problem across a wide range of domains. It is still challenging due to possession of high inter-class similarity, class imbalance datasets, and variability in data distributions. Rule-based classifiers such as XGBoost often achieve stronger performance on structured features, but they are limited in capturing smooth functional relationships among variables. Similarly, neural network models can represent complex nonlinear interactions but frequently suffer from overfitting and generalization issues. To address these limitations, we propose LFS-FRAME, a Leakage-Free Stacked ensemble framework that integrates functional learning using Kolmogorov-Arnold Networks (KAN) and rule-based learning via XGBoost for robust multiclass classification. The proposed framework constructs unbiased meta-features by employing a strict out-of-fold stacking strategy to ensure complete isolation between training and validation data hence preventing performance leakage. By learning over probabilistic outputs from heterogeneous base learners, the meta-classifier effectively exploits both global functional patterns and sharp decision boundaries present in the complex data. Experimental evaluations on multi-class datasets demonstrate that LFS-FRAME improves performance metrics, and overall accuracy is 89.85% in identifying major families and 81.74% in identifying sub-families relative to strong single-model baselines. These results highlight the effectiveness of leakage-free functional and rule-based stacking for reliable and generalizable multiclass classification.

View source

Similar papers

#machine learning Preprint Sep 2026

A Kernel-Based Modular Discriminant Analysis Framework for Small-Sample Learning

The small-sample-size (SSS) problem remains a fundamental challenge in machine learning when labeled data are scarce due to cost, accessibility, or ethical constraints. While numerous approaches have been proposed, existing methods often struggle to maintain stable and discriminative representations under high-dimensio...

Ling-Xiao Qu, Yan Pei · 1 citation
Open access Aug 2026

Heterogeneous stacking framework with adaptive Borderline-SMOTE for imbalanced binary classification: a multi-institutional validation study

Binary classification in imbalanced tabular datasets remains a significant challenge in machine learning, as conventional risk-stratification models exhibit limited discriminative performance and fail to capture nonlinear interactions among heterogeneous features. Existing approaches often suffer from three critical li...

Yang Zhang, Yanping Zhu, Xiaohui Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Meta-Learning for Classifier Selection in Image Datasets: A Feature-Driven Framework for Accuracy Prediction

No Free Lunch theorem implies that any performance gains achieved by a classifier on a particular image distribution are necessarily offset by a loss of performance over the set of all possible problems; thus, no single model is universally optimal. Selecting the most suitable classifier for image datasets is a critica...

Zahra Nabizadeh-ShahreBabak, Farzaneh Koohestani, Nader Karimi et al. · 0 citations
Open access Aug 2026

Adaptive Ensemble Learning for Accurate Classification of High-Dimensional Data

The proliferation of high-dimensional data in genomics, text analytics, hyperspectral imaging and industrial sensing has exposed a persistent weakness of conventional classifiers: as the number of features grows far beyond the number of available samples, distance measures lose contrast, decision boundaries become unst...

Porwal Rabins · 0 citations
Open access Sep 2026

Performance Enhancement on Classification of Imbalanced Data using eXtreme Gradient Boosting (XGBoost)

In real world applications, data are generated with uneven distribution called imbalance data which consists of majority and minority classes. For imbalanced data most of the classifier is biased towards the majority class. This means the classifier provides good accuracy for the majority class but very poor accuracy f...

Amit Rauniyar, Prem Chandra Roy, Anisha Pokhrel et al. · 0 citations
Open access Aug 2026

ASWBoost: Classification algorithm for noisy and imbalanced data based on parametric exponential loss

AdaBoost, a classical boosting ensemble algorithm, is widely applied for its strong classification performance. However, its standard exponential loss is highly sensitive to outliers, prone to overfitting, and inherently biased toward the majority class under class-imbalanced settings, degrading overall performance. To...

Fei Meng, Mei Yan, Hang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.