Skip to content
Preprint

DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers

Aug 2026 · 0 citations · 34 references
Computer Science Mathematics

TL;DR

Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors, significantly reduces the split search space for classification tasks while preserving predictive performance.

Abstract

Decision tree-based models are widely used in machine learning due to their interpretability and strong empirical performance. However, training decision trees can be computationally expensive, particularly for large and high-dimensional datasets, largely due to the exhaustive search over candidate splits at each node. To improve computational efficiency, we propose Data-Informed Centroid Splitting (DICS), a clustering-based framework that constructs a compact and informative set of candidate splits using data-driven priors. By incorporating class-aware structure, DICS significantly reduces the split search space for classification tasks while preserving predictive performance. We further provide theoretical analysis showing that under the stated assumptions, DICS does not degrade the performance of classification trees compared to exhaustive split search. DICS can be incorporated into classification trees, random forests, and gradient-boosting models. Extensive experiments demonstrate that DICS achieves comparable accuracy while substantially reducing training time across synthetic and benchmark datasets, highlighting the benefit of integrating data-informed priors into split selection for scalable classification tree learning.

View source

Similar papers

Aug 2026

Weighted Optimal Classification Forests

This paper introduces weighted optimal classification forests (WOCFs), a new family of classifiers that takes advantage of an optimal ensemble of decision trees to derive accurate and interpretable classifiers. We propose a novel mathematical optimization-based methodology that jointly constructs an ensemble of decisio...

Víctor Blanco, Alberto Japón, Justo Puerto et al. · 0 citations
Open access Aug 2026

Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning

These models combine the periodic splitting strategy of Hoeffding Trees, which fosters ensemble diversity, with adaptive splitting mechanisms that employ change detection algorithms to identify performance decay and determine split points.

Daniel Nowak Assis, J. P. Barddal, Fabrício Enembreck · 0 citations
Preprint Aug 2026

Diversity-Based Active Learning: An Evaluation of Metric Spaces for Active Learning Selection

Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.

Siddharth Chilamkur, D. Hochbaum · 0 citations

On the Influence of Hyperparameters in Tree-Based Linear Methods for Extreme Multi-Label Text Classification: Insights for Efficient and Effective Search

This study innovatively analyzes the hyperparameters of tree-based linear methods and suggests an efficient and effective guideline that leads to consistent improvements across datasets, thereby strengthening tree-based linear methods as a stronger XMTC baseline.

Kuan-Ting Chen, Hung-Chih Chiang, Chih-Jen Lin · 0 citations
Open access Aug 2026

Fuzzy Random Forest: Integrating Fuzzy Set Theory for Enhanced Imbalanced Classification

Standard Random Forest algorithms assume crisp class boundaries and precise feature values, limitations that become critical when dealing with ambiguous or overlapping data patterns common in imbalanced datasets. This paper presents Fuzzy Random Forest (FRF), a novel ensemble method that integrates fuzzy set theory int...

James Omusula Atsali · 0 citations
Open access Aug 2026

ASWBoost: Classification algorithm for noisy and imbalanced data based on parametric exponential loss

AdaBoost, a classical boosting ensemble algorithm, is widely applied for its strong classification performance. However, its standard exponential loss is highly sensitive to outliers, prone to overfitting, and inherently biased toward the majority class under class-imbalanced settings, degrading overall performance. To...

Fei Meng, Mei Yan, Hang Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.