A comprehensive empirical study of the impact of label-level noise on multi-label kNN classification and on Multi-label Prototype Generation (MPG) methods reveals that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful.
Abstract
Multi-label classification methods based on the k-Nearest Neighbor (kNN) rule are widely used due to their simplicity and competitive performance, but their behavior under label-level noise remains insufficiently understood, especially when combined with data reduction techniques. This paper presents a comprehensive empirical study of the impact of label-level noise on multi-label kNN classification and on Multi-label Prototype Generation (MPG) methods. We formalize six label-level noise induction policies—Additive, Subtractive, Additive-Subtractive, Distribution-Aware Additive-Subtractive, Partial Uniform Multi-label, and Swap—parameterized by both the proportion of affected instances and a severity parameter. Their effect is analyzed on three representative kNN-based multi-label classifiers (BRkNN, LPkNN, and MLkNN) and five MPG strategies (MRHC, MChen, MRSP1-3) across eight benchmark datasets with varying label cardinality and imbalance, comprising an extensive experimental grid of 816,480 configurations. The results reveal that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful. Moderate neighborhood sizes (around \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k=7$$\end{document}) provide a good trade-off between robustness and accuracy, while MLkNN is consistently the most resilient classifier under severe noise. Among MPG methods, MRSP3 emerges as the most robust reduction strategy, whereas aggressive reductions, particularly with MRHC, can amplify the negative effects of noise. The code and complete experimental results are publicly released to support reproducibility and further research.
Multi-label data often contain high-dimensional features, outlier instances, and noisy labels, all of which can lead to the curse of dimensionality and decreased performance in downstream tasks. Although numerous data reduction methods have been developed, existing approaches face two major limitations: 1) existing met...
Li Yang, Yan-Yong Huang, Jin-Yuan Chang et al.· Proceedings of the Thirty-Fi...· 0 citations
Distance-based multi-label classifiers that rely on a fixed, non-adaptive metric—ML-kNN, RBF-kernel machines with fixed bandwidth, and analogous templates—compare instances through symmetric feature differences and therefore do not encode whether label evidence lies above or below a typical feature value; adaptive metr...
Filippo Casu, A. Lagorio, G. Trunfio· Data mining and knowledge di...· 0 citations
This study innovatively analyzes the hyperparameters of tree-based linear methods and suggests an efficient and effective guideline that leads to consistent improvements across datasets, thereby strengthening tree-based linear methods as a stronger XMTC baseline.
Tri-training is a classic semi-supervised learning framework that improves classifier performance by exploiting unlabeled data. However, it suffers from invalid view redundancy assumption and severe pseudo-label noise in real-world applications, which leads to performance degradation. To address these problems, this pa...
Xiangxiang Cai, Song Li, Yulin Zhang· International Conference on...· 0 citations
This work proposes a novel hybrid approach that combines the strengths of the extreme learning machine (ELM) and the twin support vector machine (TSVM) to address the challenges of robustness and scalability in multi-label classification, particularly in settings where deep learning is not practical due to limited trai...
Amisha Bharti, Vasudha Bhatnagar, Vikas Kumar· International Journal of Dat...· 0 citations
Experiments on two multi-expert ulcerative colitis endoscopic-image datasets under two ordinal-noise models show that Ord-NLL is competitive with or superior to strong baselines while reducing mean absolute error, and that Ord-NLL+ often yields further gains.
Shumpei Takezaki, K. Shiku, S. Harada et al.· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.