A novel approach that leverages statistical redistribution in the output space to approximate the post-removal confidence vectors of a retrained model, alleviating scalability limitations and potentially mitigates privacy concerns inherent to data-dependent solutions is proposed.
Abstract
Label removal occurs frequently in classification systems with evolving taxonomies, where categories must be dynamically updated or eliminated. To accommodate such changes, classification models must adapt accordingly. Existing solutions, broadly categorized as retraining-based and feature-space-adjustment-based, share common limitations despite their variations, including reliance on access to original data, substantial computational and storage costs, inconsistent results, poor scalability, and degradation of model utility. To address this, we propose a novel approach that leverages statistical redistribution in the output space to approximate the post-removal confidence vectors of a retrained model. Applicable as a modular output filter, our method bypasses the burden of feature-space adjustments or loss-function convergence, alleviating scalability limitations. Furthermore, by requiring only existing labels and prior output confidences, the method potentially mitigates privacy concerns inherent to data-dependent solutions. Extensive experiments demonstrate competitive performance against full retraining, with improvements in computational efficiency and privacy preservation across several classification tasks.
The MLIDSC introduces an automated labeling mechanism that eliminates the need for prior parameter assumptions and uses a novel weighted scheme that combines the imbalance ratio and the importance of individual instances at a given time, ensuring a focus on critical data points.
Bohnishikhan Halder, K. M. Azharul Hasan, Md. Manjur Ahmed· Applied intelligence (Boston...· 0 citations
This work proposes distilling the model into a probabilistic classifier, enabling lightweight deployment without repeated LLM calls, and demonstrates that LSR improves macro-F1 scores by an average of 7.0% compared to standard zero-shot classification baselines.
Nathan Vandemoortele, Bram Steenwinckel, F. Ongenae et al.· Discover Computing· 0 citations
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label noise is critical fo...
Faizaan Ali, Inwon Kang, O. Seneviratne· 0 citations
A novel oversampling algorithm: the adaptive weighting–synthetic minority oversampling technique (AW-SMOTE), which combines the two perspectives of boundary tightness and local density and provides global sample enhancement support.
This paper proposes ADA-CS, a plug-and-play module compatible with any ADA or ASFDA framework, and introduces a CSS metric to quantify the Concept Shift Severity across domains, revealing that non-negligible concept shift exists in many transfer tasks.
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026