Jul 2026· IEEE Transactions on Neural Networks and Learning Systems· Vol PP, pp. 1-15· 0 citations
Medicine
TL;DR
A novel weakly supervised (WS) learning MLTC framework consisting of a novel category word selection method, namely category word selection with significance ranking and crowd-sourcing (Cws-src), and a generic WS learning MLTC method, namely WS multilabel text classification with correlation-aware label propagation (Wmltc-clp), which estimates accurate pseudolabels by propagating them over a text correlation graph.
Abstract
Multilabel text classification (MLTC) methods require enormous labeled training samples to ensure the model's performance, which involves significant manual labor costs. An alternative to conducting MLTC is to only employ predefined representative words of classes, namely category words, as the weak supervision. In this article, we propose a novel weakly supervised (WS) learning MLTC framework consisting of two parts. First, we propose a novel category word selection method, namely category word selection with significance ranking and crowd-sourcing (Cws-src), which generates confident category words by manually selecting from the topically reranked words using a new TW-ITF weighting scheme, thereby effectively mitigating the noises in pseudolabels by filtering repetitive and less significant terms for each class, leading to improved classification performance. Subsequently, we propose a generic WS learning MLTC method, namely WS multilabel text classification with correlation-aware label propagation (Wmltc-clp), which estimates accurate pseudolabels by propagating them over a text correlation graph. To evaluate the proposed framework, we conduct extensive experiments on nine benchmark datasets, including five sentiment analysis datasets and four prevalent MLTC datasets. The results demonstrate that Cws-src can generate more confident category words and Wmltc-clp can achieve significant improvements over the WS learning baselines. The maximum performance gains of Wmltc-clp over the best WS learning baseline methods reach 0.096, 0.081, 0.075, and 0.02 on Micro- $F1$ , Macro- $F1$ , average precision (AP), and ranking loss (RL) across all benchmark datasets.
Multi-label classification of financial news is frequently affected by incomplete and noisy annotations, while obtaining expert-curated labels at scale is prohibitively expensive. This study proposes a weakly supervised classification framework that combines large language model (LLM) zero-shot annotation with a serial...
A framework is proposed that identifies which label pairs the model struggles to distinguish, expands the candidate set to include confusable labels, and generates targeted rules to differentiate between similar candidates, which requires no fine-tuning and transfers to smaller, cheaper models.
Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each...
This work proposes distilling the model into a probabilistic classifier, enabling lightweight deployment without repeated LLM calls, and demonstrates that LSR improves macro-F1 scores by an average of 7.0% compared to standard zero-shot classification baselines.
Nathan Vandemoortele, Bram Steenwinckel, F. Ongenae et al.· Discover Computing· 0 citations
This paper proposes a novel Partial label-based Self-training framework (PaSta) that leverages partial label learning technique to overcome the limitations of existing methods and designs a partial label-based classification model with two well-crafted loss functions to guide the model learning at both label and repres...
Yujing Liu, Yi-Xin Liu, Yu Zheng et al.· 0 citations
Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change...
Ebenezer Tarubinga· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.