Skip to content
Open access

Exploring the impact of label-level noise on multi-label k-Nearest Neighbor classification

Jul 2026 · Pattern Analysis and Applications · Vol 29 · 0 citations · 33 references
Computer Science

TL;DR

A comprehensive empirical study of the impact of label-level noise on multi-label kNN classification and on Multi-label Prototype Generation (MPG) methods reveals that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful.

Abstract

Multi-label classification methods based on the k-Nearest Neighbor (kNN) rule are widely used due to their simplicity and competitive performance, but their behavior under label-level noise remains insufficiently understood, especially when combined with data reduction techniques. This paper presents a comprehensive empirical study of the impact of label-level noise on multi-label kNN classification and on Multi-label Prototype Generation (MPG) methods. We formalize six label-level noise induction policies—Additive, Subtractive, Additive-Subtractive, Distribution-Aware Additive-Subtractive, Partial Uniform Multi-label, and Swap—parameterized by both the proportion of affected instances and a severity parameter. Their effect is analyzed on three representative kNN-based multi-label classifiers (BRkNN, LPkNN, and MLkNN) and five MPG strategies (MRHC, MChen, MRSP1-3) across eight benchmark datasets with varying label cardinality and imbalance, comprising an extensive experimental grid of 816,480 configurations. The results reveal that Additive and Partial Uniform noise are the most detrimental, whereas Subtractive and cardinality-preserving policies are comparatively less harmful. Moderate neighborhood sizes (around \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$k=7$$\end{document}) provide a good trade-off between robustness and accuracy, while MLkNN is consistently the most resilient classifier under severe noise. Among MPG methods, MRSP3 emerges as the most robust reduction strategy, whereas aggressive reductions, particularly with MRHC, can amplify the negative effects of noise. The code and complete experimental results are publicly released to support reproducibility and further research.

Read PDF

Similar papers

Conference Open access Sep 2026

EMMS: Evidential Multi-Label Multi-Dimensional Selection

Multi-label data often contain high-dimensional features, outlier instances, and noisy labels, all of which can lead to the curse of dimensionality and decreased performance in downstream tasks. Although numerous data reduction methods have been developed, existing approaches face two major limitations: 1) existing met...

Li Yang, Yan-Yong Huang, Jin-Yuan Chang et al. · 0 citations
Open access Sep 2026

Direction-aware multi-label feature selection via paired signed-deviation lifting

Distance-based multi-label classifiers that rely on a fixed, non-adaptive metric—ML-kNN, RBF-kernel machines with fixed bandwidth, and analogous templates—compare instances through symmetric feature differences and therefore do not encode whether label evidence lies above or below a typical feature value; adaptive metr...

Filippo Casu, A. Lagorio, G. Trunfio · 0 citations

On the Influence of Hyperparameters in Tree-Based Linear Methods for Extreme Multi-Label Text Classification: Insights for Efficient and Effective Search

This study innovatively analyzes the hyperparameters of tree-based linear methods and suggests an efficient and effective guideline that leads to consistent improvements across datasets, thereby strengthening tree-based linear methods as a stronger XMTC baseline.

Kuan-Ting Chen, Hung-Chih Chiang, Chih-Jen Lin · 0 citations
Conference Aug 2026

Improved tri-training semi-supervised classification algorithm based on adaptive neighborhood entropy

Tri-training is a classic semi-supervised learning framework that improves classifier performance by exploiting unlabeled data. However, it suffers from invalid view redundancy assumption and severe pseudo-label noise in real-world applications, which leads to performance degradation. To address these problems, this pa...

Xiangxiang Cai, Song Li, Yulin Zhang · 0 citations
Aug 2026

Multi-label classification with extreme learning machine and twin support vector machine: a novel hybrid framework

This work proposes a novel hybrid approach that combines the strengths of the extreme learning machine (ELM) and the twin support vector machine (TSVM) to address the challenges of robustness and scalability in multi-label classification, particularly in settings where deep learning is not practical due to limited trai...

Amisha Bharti, Vasudha Bhatnagar, Vikas Kumar · 0 citations
Open access 2026

Ord-NLL: Negative Label Learning for Ordinal Noisy Labels

Experiments on two multi-expert ulcerative colitis endoscopic-image datasets under two ordinal-noise models show that Ord-NLL is competitive with or superior to strong baselines while reducing mean absolute error, and that Ord-NLL+ often yields further gains.

Shumpei Takezaki, K. Shiku, S. Harada et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.