Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.
Abstract
Decision tree-based models are widely used in machine learning due to their interpretability and strong empirical performance. However, training decision trees can be computationally expensive, particularly for large and high-dimensional datasets, largely due to the exhaustive search over candidate splits at each node. To improve computational efficiency, we propose Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors. By incorporating class-aware structure, DICS significantly reduces the split search space for classification tasks while preserving predictive performance. We further provide theoretical analysis showing that under the stated assumptions, DICS does not degrade the performance of classification trees compared to exhaustive split search. DICS can be incorporated into classification trees, random forests, and gradient-boosting models. Extensive experiments demonstrate that DICS achieves comparable accuracy while substantially reducing training time across synthetic and benchmark datasets, highlighting the benefit of integrating data-informed priors into split selection for scalable classification tree learning.
This paper introduces weighted optimal classification forests (WOCFs), a new family of classifiers that takes advantage of an optimal ensemble of decision trees to derive accurate and interpretable classifiers. We propose a novel mathematical optimization-based methodology that jointly constructs an ensemble of decisio...
Víctor Blanco, Alberto Japón, Justo Puerto et al.· INFORMS journal on computing· 0 citations
These models combine the periodic splitting strategy of Hoeffding Trees, which fosters ensemble diversity, with adaptive splitting mechanisms that employ change detection algorithms to identify performance decay and determine split points.
Daniel Nowak Assis, J. P. Barddal, Fabrício Enembreck· Data mining and knowledge di...· 0 citations
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
This study innovatively analyzes the hyperparameters of tree-based linear methods and suggests an efficient and effective guideline that leads to consistent improvements across datasets, thereby strengthening tree-based linear methods as a stronger XMTC baseline.
Standard Random Forest algorithms assume crisp class boundaries and precise feature values, limitations that become critical when dealing with ambiguous or overlapping data patterns common in imbalanced datasets. This paper presents Fuzzy Random Forest (FRF), a novel ensemble method that integrates fuzzy set theory int...
James Omusula Atsali· Asian Journal of Probability...· 0 citations
AdaBoost, a classical boosting ensemble algorithm, is widely applied for its strong classification performance. However, its standard exponential loss is highly sensitive to outliers, prone to overfitting, and inherently biased toward the majority class under class-imbalanced settings, degrading overall performance. To...
Fei Meng, Mei Yan, Hang Liu et al.· PLoS ONE· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.