Ensemble–based instance selection for multi–target regression using kernel sparse modeling
Abstract
This paper proposes a novel instance selection (IS) method for multi-target regression (MTR). The proposed method introduces Sparse Modeling Representative Selection (SMRS) to characterize the self-representativeness of instances, where a block-sparse coefficient matrix is learned to identify representative training samples. Considering that MTR data contain both input-space structures and multi-target output-space dependencies, we assume that the self-representative relationships in the input and output spaces should be consistent. Based on this assumption, a Frobenius-norm-based discrepancy minimization term and a Hilbert-Schmidt Independence Criterion (HSIC)-based relevance maximization term are incorporated to encourage the model to learn consistent and statistically dependent self-representation structures across the two spaces. To capture nonlinear instance relationships, the proposed model is further extended to a kernelized formulation. In addition, an ensemble framework is developed to exploit target correlations from multiple target-aware subspaces, thereby improving the robustness of instance selection. A parameter-free adaptive thresholding strategy is also designed to automatically determine the selected subset from the learned instance-score distribution. Extensive experiments on multiple benchmark MTR datasets demonstrate that the proposed methods can effectively reduce the training data size while preserving or improving downstream predictive performance compared with representative instance selection and data reduction baselines.